Showing posts with label searching speech. Show all posts
Showing posts with label searching speech. Show all posts

Friday, November 19, 2010

Search your own dogfood

How many hours do I spend writing deliverables and reports? I'd rather not count. Here I am on Friday night with a to do list left over from the week that seems only very vaguely connected to my main mission as a researcher, namely to improve multimedia access systems, especially for spoken audio and video with a speech track.

Sometimes it takes writing a blog entry to refocus on the core values of multimedia search. I was going through the pictures from the Searching Spontaneous Conversation Speech workshop in order to find a good one to add to the latest newsletter report, and dang it, if there weren't so many speaker pictures that we ruined because Florian is crouching in the middle in the front, tending to the laptop that we were using to capture the sound.

At Interspeech we discussed the idea of simply recording all the spoken audio at both the MediaEval 2010 workshop and the SSCS 2010 workshop in order to start an audio corpus of workshops to use for research on meeting retrieval. It sounded like a good idea that we would never have the time to pull off, but sure enough, there we were in Italy, and a network of people came together and brought sound equipment from all over and we had ourselves a system for audio capture. I remember the satisfaction in his voice, when Florian announced "We are now recording six channels". Actually, I remember it because I listened to it on the recording afterwards as we started the laborious process of post-processing and I wondered "Gee, what kinds of things were we talking about next to the main presentations."

So here's the refocus. Florian isn't actually ruining the picture. His presence actually underlines what the speaker is talking about -- the slide reads "The ACLD: Speech-based Just-in-Time Retrieval of Meeting Transcripts, Documents and Websites". We have made such a huge step in this direction that in are own lives we can simply decide to capture our spoken content, everyone at the workshop says, "OK, that's cool" and bang we have more data than we know what to do with.

We also did this at SSCS 2008 in Singapore. The videos were online for a while -- we transcribed them using Nuance Audiomining SDK for speech recognition and made them searable with a Lemur-based earch engine. For awhile, we could visit a website and search our own dogfood, as it were. It seems, however, that the multimedia lifecycle got the better of our content: the system was not maintained and now the videos are no longer available online. I don't know if we'll do much better this year, but the point is that we keep on trying. And we have Florian in the middle of the workshop picture reminding us that this attempt may be time consuming, but it is constitutes the core of our research mission.

Friday, October 29, 2010

ACM Multimedia SSCS 2010 Workshop on Searching Spontaneous Conversational Speech

The Fourth Workshop on Searching Spontaneous Conversational Speech took place on 29 October 2010 at ACM Multimedia. Papers were presented about techniques for speech retrieval, speaker role recognition, spoken term detection and concept detection. Invited speakers addressed challenges for the future of spoken content retrieval, including interview data, multimedia archives and the Spoken Web. The demonstrations were a highlight of the workshop. These were first introduced in a boaster session and then presented to workshop participants in an interactive session. Here's the Wordle Word cloud made from the title and the abstracts of all the papers presented!

Currently, we are getting ready for an upcoming special issue on searching speech in ACM Transactions on Information Systems.

Saturday, April 10, 2010

Speechless? Not us.


Our proposal for a fourth workshop on Searching Spontaneous Conversational Speech at ACM Multimedia 2010 was accepted today.

ACM Multimedia 2010 Workshop
Searching Spontaneous Conversational Speech (SSCS 2010)
29 October 2010, Firenze Italy

http://www.searchingspeech.org/

The SSCS 2010 workshop is a forum for presentation of recent research results concerning advances and innovation in the area of spoken content retrieval and in the area of multimedia search that makes use of automatic speech recognition technology. Spontaneous, conversational speech occurs in a wide variety of domains and the workshop is relevant for lectures, meetings, interviews, debates, conversational broadcast (e.g., talkshows), podcasts, call center recordings, cultural heritage archives, social video on the Web and spoken natural language queries. The objective of the workshop is to bring together researchers in theareas of speech recognition, audio processing, multimedia analysis and information retrieval for exchange and interaction.

Friday, October 23, 2009

SSCS 2009


The Third Workshop on Searching Spontaneous Conversational Speech (SSCS 2009) took place on 23 October 2009 in Beijing China in conjunction with ACM MultiMedia 2009. We had a great set of demos and talks. As an organizer this gives you a warm pleased feeling -- all that work is worth it. Domains covered included broadcast, meetings, interviews, telephone conversations, podcasts and voice tagging for photos. The approaches presented involved using a variety of techniques including subword units, exploiting dialogue structure, fusing retrieval models, modeling topics and integrating visual features. Such events serve to highlight the importance of the spoken word in many multimedia access and retrieval applications. And also to remind us how far we are from exploiting it fully.

Thursday, September 10, 2009

Getting the words right

The final day of Interspeech 2009 here in Brighton. It's been a great conference and each and every keynote has been well worth getting up for. This morning, Mari Ostendorf talked about "Transcribing Speech for Spoken Language Processing." Interspeech encompasses a staggeringly broad spectrum of perspectives on speech research and technology. For every point here, there is an immediate counterpoint, and it was without doubt under influence of this chorus that the opening slide of the keynote this morning displayed a long-play version of the title reminding the audience that they would be hearing about transcribing human-directed human speech, as opposed to speech that humans produce to communicate with computers.

The message from the keynote that will ring longest in my ears was, "The goal of speech transcription is information access." This leaves open of the course, the question of what is the information and what is the access when it comes to content that contains the spoken word. I find myself compiling little lists of domains in which information encoded in spoken audio could be important: podcasts, video diaries, lifelogs, meetings, call center recordings, social video networks, Web TV, conversational broadcast, lectures, discussions, debates, interviews and cultural heritage archives, home videos, photo annotations, video conferences. These lists invariable end with etc. etc. etc. And what constitutes access (keyword search, retrieval, question answering, browsing, recommendation...) is another question to which we can't give a closed-set answer.

My personal experience doesn't really support the idea that we need to push the envelope. The last video I watched I found because a link was sent to me by my cousin. The content of the video was a short clip of her new cat purring. No real access problem there. No information either. The purr did not inform me in the conventional sense. In fact, there wasn't much human speech involved at all. Nonetheless, I found the content supremely worthwhile of my watching time. Although my own multimedia access needs are a string of examples of this pre-solved sort, I do agree that the challenge of access to speech-based information is a serious one and will require a great deal of effort to address.

The full phrase in Ostendorf's slide read, "The goal of speech transcription is information access, not just getting the words right." But maybe it is about "getting the words right". The words referred to are, presumably, white-space delineated grapheme strings, lexical words, citation forms. But we can also see a word as the totality of knowledge that a human needs to possess in order to deploy it in human-to-human communication. There may be a limit to how far we can go beyond that sort of word and still remain within what is meaningful in the context of our information access needs.

We can go for prosody, for speech act, subjectivity, affect, but in the end we'll never capture the "you had to have been there" component of understanding. And already the moment of that particular purr video has passed and my next need for video content will be for a new one.

Friday, May 8, 2009

WIAMIS 2009 in London

Attended the International Workshop on Image Analysis for Multimedia Interactive Services a.k.a. WIAMIS 2009. The PetaMedia task force organized a special session to showcase topics related to combining multimedia content analysis with user contributed information and social network structure. We presented some initial work on combining speech-based indexing features with low level visual information for improved video retrieval.

Thursday, April 9, 2009

ECIR 2009 in Toulouse


I've never attended the European Conference on Information Retireval before, but this year I flew to Toulouse for ECIR 2009. My continued reflection on the larger implications of the properties of the error produced by speech recognition systems finally yielded some fruit...but it is still only a small window on a larger story.

Larson M., Tsagkias E., He J., de Rijke M., Exploring the Global Semantic Impact of Speech Recognition Errors on Spoken Content Retrieval, 31st European Conference on Information Retrieval Conference (ECIR 2009).

Tsagkias E., Larson M., de Rijke M., Exploiting Surface Features for the Prediction of Podcast Preference, 31st European Conference on Information Retrieval Conference (ECIR 2009).