Monday, July 26, 2010

Advances in Multimedia Retrieval Tutorial at ACM Multimedia 2010

I worked a long day today and now I'm home and have eaten dinner and thinking about how to relax. I'd like to watch an episode of Merlin on the Internet. Preferably legal and one I've never seen before -- wouldn't mind paying if it was a site I trusted. That seems like a pretty complicated search and my prediction is that it will lead to frustration. So here I am writing in my blog instead.

Today the longplay description of our upcoming tutorial on Frontiers in Multimedia Search went online. We want to start out by addressing the question of how can multimedia search benefit people's daily lives, at work and otherwise. I'm feeling rather a strong need for the benefit of multimedia at the moment. If I can't have my Merlin, right now I wouldn't mind browsing back through recordings of the SIGIR presentations that I heard last week -- and maybe some of the ones that I have missed.

Then we are planning to go on to take a look at new approaches to multimedia retrieval that we divide into three categories (I include a couple of my own notes on each):
  • Making the most of the user: In motion on the Internet we dribble information behind us. We tag, we query, we click, we brush over a page without a second glance. We have the capacity to glance at a set of snippets, glaze over what does not interest us to find what does. Making the most of the user is about letting the search engine turn the computational crank and do the look ups, leaving those fine-grained semantic judgments to the human brain.
  • Making the most of the collection: Sometimes the collection can speak for itself. Pseudo-relevance feedback may dilute our queries, but it also is a valuable tool for increasing recall. And then there is collaborative filtering: Making use of patterns that we as users leave behind -- but now at the collection or community level.
  • Making the most of individual items: What is important here is how you can do the best that you can with noisy sources of features (speech recognition, visual concept detection) to represent items. You don't need to necessarily provide a complete representation of an item -- information that helps distinguish items or keeps them from getting confused can sometimes be a big help.
All in all, the aim is to present our favorite new multimedia search work -- working to inject new techniques and perspectives from the IR community and speech community into multimedia. And, of course, develop our own understanding of multimedia search along the way. Maybe I will then also feel better equipped to find something that I could watch now that would make me as happy as Merlin. Or at least to understand exactly what is still missing...

Friday, July 23, 2010

SIGIR 2010 Crowdsourcing for Search Evaluation Workshop

We used Amazon Mechanical Turk (MTurk) to gather annotations for the video corpus to be used in the Affect Task at the MediaEval 2010 benchmark evaluation. The task involves automatically identifying videos that viewers report to be particularly boring. We wrote the corpus development up and submitted it to the Crowdsourcing for Search Evaluation Workshop at SIGIR 2010. Initially we wondered a bit if the paper was appropriate for the workshop, since we were working on affect and not directly on search. But we were glad that we took the risk and went for it. The paper was accepted and the workshop was great -- right on target with our interests.

Soleymani, M. and Larson, M. Crowdsourcing for Affective Annotation of Video: Development of a Viewer-reported Boredom Corpus.
In Proceedings of the SIGIR 2010 Workshop on Crowdsourcing for Search Evaluation.

We also received the Runner Up for the Most Innovative Paper Award, which was sponsored by Microsoft Bing. Thank you! We are already considering how to get the most bang for our Bing bucks. Probably it will flow directly back into MTurk for our next crowdsourcing project.

Sunday, July 18, 2010

Emotive speech and navigation systems

During a recent family weekend, my best friend, my mother and I found ourselves in a car using my aunt's navigation system to guide us to our destination. We quickly developed a love-hate relationship with the device -- our feelings of annoyance generally outweighing our gratefulness of having been guided efficiently to our destination.

My mom then circulated an article from CNN entitled Why GPS voices are so condescending. And my aunt mailed back, 'Hey, isn't this what you work on?'

The answer to that question is, yes, well, not quite. I go the other direction. Instead of automatically producing emotive speech, I start with a recording of emotive speech and automatically analyze how it was produced. We just got a paper accepted in a session entitled "Paralanguage" at Interspeech 2010 :

Jochems, B., Larson, M., Ordelman, R., Poppe, R. and Thruong, K. Towards Affective State Modeling in Narrative and Conversational Settings. Proceedings of Interspeech 2010 (to appear).

The CNN article also falls into the category of paralanguage. Paralanguage is basically the things that we do with speech that modifies the factual content or conventional meaning of what we are saying. In this case, it's adding emotive nuance.

The designers of navigation systems are stuck with the following impasse: If a navigation system is good, it will always be right. Socially, there is a tabu against always being right. An always-right system will always be perceived as condescending, be its voice ever so loving and sweet. That's simply the way that social behavior works -- we count on each other to act responsibly, but not to pretend that we're perfect. The implication is that we, as humans, will never truly adopt the metaphor of "it's just a person telling me where to drive" for a navigation system if that system's understood purpose is to deliver infallibility.

In my opinion, what the designers of navigational systems should do is to use the voice of someone who enjoys special social status and as such, "gets away" with being always right. For example, theoretical physicist Stephen Hawking. His smarts are generally acknowledged to transcend the smarts of the rest of us mere mortals. Interestingly, he also speaks using a computer voice because he has neuro-muscular distrophy. It wouldn't take a whole lot of memory space on your little navigation device in order to produce a believable rendition of his speech.

The issue also has a huge safety aspect (which is also raised in the CNN article). If the navigational system uses emotive speech in a very convincing manner, it is smooth sailing. However, what if something goes wrong? To the driver, it will be like a thunder-bolt out of a blue sky. Everything was going fine, and all of a sudden the device turned and lashed-out with an emotively inappropriate direction. Possibly, this would happen at a critical driving point. The driver shouldn't be so comfortable with the device as to completely exclude the possibility that it goes way off the mark.

Basically, a car navigation system presents us with another instance of the Paradox of 'Simplicity'. It takes a lot of very complicated innards to make a device that drivers perceive simply as a human telling us how to get there. The paradox comes in when that device does something wrong and all of a sudden the human is stuck both solving the immediate driving issue and also compensating for the apparently inexplicable (those complicated innards!) failure of the system. In this case, for example, a beautifully real rendition of a plaintive tone pleading "Turn back! Turn back!" when actually we find ourselves stuck in the express lane in heavy traffic.

The theoretical physicist persona would help to lessen the impact of such errors. Sorry, Stephen Hawking, but theoretical physicists can get away with being socially inappropriate once in a while without throwing us into a state of shock -- we assume that they are simply busy on a higher plane and don't mean to really insult or confuse us.

However, instead of talking to Stephen Hawking about a deal to have him donate his authority to make navigational systems safer, navigational system companies (according to CNN) are looking into fitting the systems with the driver's own voices. It sounds cool, until you think about some of the implications.

First, there are probably people who don't react well to their own voices. Perhaps I could accept my own voice reminding myself of the route to somewhere I've been before, but my own voice directing me to somewhere I have never been, for example, Makuhari, Japan (where Interspeech 2010 will take place in September) is absolutely implausible. I know I can't trust myself on that one.

Second, drivers need to be encouraged not to turn off their human intelligence when driving with a navigational system. The system doesn't tell you, "Stop here, the light is red". Listening to your own voice is probably not the right way to ensure that you are actively applying the underlying rules and your own common sense to driving.

Third, it's not uncommon to rent a car borrow someone else's car or navigational system on a single-case basis. For example, my aunt lent us her device for one trip. Wouldn't we like our devices more if they were one-size-fits all? Just as Walter Cronkite provided widespread satisfaction as the voice of the evening news, what's wrong with generally-acceptable central voice for all navigation systems?

Fourth, it's not only the driver would needs to listen to the navigation system. With several people in the car, navigation often involves pooling knowledge of the route and negotiating consensus. If the driver's voice is talking on the navigational system, the passengers are shut out of the process. For maximally safe driving, you don't want a "back seat driver", but a co-pilot who is engaged in the process is very helpful.

Fifth, it is not clear that the navigation system companies are the ones that should be making the decision about how navigation system personae can be made more acceptable to drivers. If they can convince individual drivers that they need to have a personalized voice for their system will open up an incredible new opportunity for profit for navigation device companies. On top of the system and the route information, they will also be able to sell you your own personna.

Additionally, a universal "Stephen Hawking" solution, which I am arguing may actually be safer, would make it impossible for navigation system companies to distinguish themselves from each other on the basis of the differential appeal of their navigation personna and is simply not in companies best business interest.

My suggestion is simply to learn to love the condescending dead-pan delivery of your current navigation system -- demanding anything different may prompt the designers of navigation systems to make the situation a whole lot worse.

Don't we do this already? How often have you ever been directed somewhere by a fellow human issuing emotively inappropriate directions? You've reminded yourself to take some deep breaths, stay concentrated on the road and gotten there in the end. We shouldn't demand from our automatic devices more than what we get from our fellow human beings.

P.S. Whoa, this claims to be a blog on the topic of search, what does this have to do with search? OK. You've caught me. Sometimes I just write things here because I know that I can find them again.

Saturday, July 17, 2010

IF discouraged THEN write good reviews

Ever get a Bad Review? I don't mean one where the reviewer gives constructive criticism and recommends rejection. I mean one that is really bad in the sense that it is unhelpful, off-topic, lacking in rigor, poorly written, pedantic or pompous. It takes a lot of energy to sort through these sorts of reviews, find the wheat discard the chaff and make sure that the experience doesn't drag you down to the point of derailing a potentially productive scientific endeavor. Bad Reviews sometimes even recommend acceptance. Acceptance leads, perhaps not to disappointment at the moment, but rather to more general scientific disheartenment: Is this really the level of intellectual standards that characterizes the field to which I have chosen to devote my career?

A surprisingly satisfying way to push back against Bad Reviews, is to engage in reflection upon one's own reviewing skills and strive to improve them. This course of action is not going to have an immediately measurable effect of improving the system as a whole, but it does restore a sense of balance. Especially if you interact with a lot of students, you have an amazing opportunity to teach them to review. There's something cheering about knowing that scientists that you have mentored are not going to be the ones generating the Bad Reviews of the next generation.

In order to be able to tell people quickly about my own reviewing values and my campaign to constantly improve my own reviews, I have packed the points I consider while reviewing into a scheme that I call IF THEN:
  • I is for Issue: Does the paper motivate the issue that it addresses and then close the loop in the end, convincing the reader that it has accomplished what it set out to do?
  • F is for Fit: Does the paper fit with the call for papers of the conference or scope of the journal to which it was submitted?
  • T is for Technical soundness: Do the authors apply solid, state-of-the-art experimental and/or analytical method?
  • H is for Historical context: Do the authors present the context of their work? (Including both the related work and the outlook onto future work.)
  • E is for Exposition: Is the paper clearly written and a pleasure to read? Is the information it contains complete and comprehensive?
  • N is for Novelty: Is the idea new in the field? Is it the sort of innovation that is destined to make an impact?
If you're doing IF THEN right it should be a bit uncomfortable. Particularly difficult is the self-reflection necessary to make an honest estimate of one's own expertise in some subjects. I didn't promise this wouldn't hurt, what I promised is that constant work on your own reviewing standards eases the pain (and and controls the damage) caused by receiving a Bad Review.

But what if you are already a world class reviewer? What then? Like musicians we must remember that excellence is not static. Moshe Vardi introduced a rule for reviewing an editorial in the current edition of Communications of the ACM. The rule reads, "Write a review as if you are writing it to yourself." He calls it The Golden Rule of Reviewing. Most people are still working on putting in to practiced the Golden Rule they learned in their childhoods. The Bad Reviews will keep on coming, and about the only thing we have control over is how we review back.

Sunday, June 6, 2010

Looking for a Scientific Programmer

I'm starting a cool new project on speech-based access to images, but I am in need of a programmer. I'm trying to find the right person -- ideally I would like someone who also had an interest in the process of design and evaluation of the system. The person probably just finished their masters and is trying to get an idea of an area for a PhD, or just generally thinks that one year of experience in the Multimedia Information Retrieval Lab at Delft University of Technology would be enriching. Unfortunately, trying to get someone like this just led to me loosing my dream candidate to a PhD program in Groningen. Yikes! The project's starting on 1 September 2010.

The formal qualifications are listed below. Thanks in advance if you help me find a match between my needs and a candidate.

Profile for scientific programmer in the Delft Multimedia Information Retrieval Lab at the TU-Delft
Contact: Martha Larson m.a.larson@tudelft.nl
  • Experience developing web applications using a web development stack, (one of LAMP, Java/Tomcat, ASP.NET/C#)
  • Experience with designing HTTP-based server APIs
  • Alternatively or in addition: Experience with html/Javascript/AJAX/Flash/Silverlight
  • Alternatively or in addition: Interest or experience with Android
  • Experience with speech recognition, dialog systems, audio spatialization a benefit
  • Experience with programming in a research environment a plus
  • Proficiency in English, both spoken and written
The person needs to be an EU citizen, but if I find a great candidate who is not, I am willing to attempt to "battle the system" to get him/her.

P.S. This post represents an experiment in making use of my social network. If was were more up-to-date I suppose I'd be using Linked-In or Facebook, but neither are come naturally to me, somehow.

Friday, May 7, 2010

Relevant to the query "List of Internet Video Genres"

From the perspective of automatic multimedia content analysis, there is vast difference in visual content between video that was produced for the purpose of being understood and absorbed by an audience and video that was captured with no explicit communicative intent. If two people shake hands in a film, the audience will know that they are shaking hands and it will fit with the dialogue in the sound track and with the over all story. If two people shake hands on a surveillance video one can occlude the other or perhaps they're passing a cigarette lighter who knows.

It seems like creator intent is an important clue for visual indexing of video for retrieval. But the example above is simplistic. If we want to find video on the internet, there are a whole range of intents between film and surveilance. What are these varieties? Today I thought I could type "list of internet video genres" into my favorite mainstream search engine and have it spit me out a list of things that we as Internet users do when we make video. That didn't happen. I spend some time in amused pondering over the juxtaposition in The six most baffling genres on YouTube. But I wanted something a bit more comprehensive than that (with a different tone), so I'm posting my own list.

Captured video: Walking through the room and not knowing the camera is on. Interesting as a curiosity, but serves no specific purpose.
Life-log: I know the camera is on but don’t really think about it. Serves the purposes of off-line memory.


Surveillance video: I put the camera in some particular place to capture the scene, but the people in the scene don’t pay attention to it. Serves the purpose or providing extra ears and eyes.


Home video: I am obviously holding a camera and pointing at people. Not trying to do anything but “get the feel of the situation on video” Serves the purposes of off-line memory. The act of making the video is also inherently entertaining and it might not necessarily be watched.


Event: I am documenting an event, like a wedding. The video is meant to portray both the compliance of the event to social convention and also its uniqueness. Ideally, I want to see the face of the bride and groom and hear the “I do”. Serves the purpose of memory, but may also be considered a public document that attests to status.


Meeting video/lecture video: I know the camera is on but don’t really think about it. Serves the purposes of off-line memory and possibly institutional record. Used to rewatch things that might have been missed the first time. The camera has a specific position – other material such as slides or white board shots might be present.


Testimony: A narrator recounts and experience. Spoken audio is unscripted but declarative factual statements. The narrative is usually temporally organized. The visuals are “convenience visuals”, but there are typical camera angles: frontal shot, shot of interviewer in dialogue with the interviewee. Viewer acquires some declarative knowledge, but basically, it’s an impression of the situation. 


How-to video: Demonstrates how to do something. The visuals are key, with the camera angle chosen to give maximal information. Items depicted in the visual tract have a high probability of being named in the speech track. After watching this video the viewer is intended to have procedural knowledge of the task.


Learning video: The video acts out scenes and viewers are invited to put themselves into the scene. This video is a surrogate for experience. Here, the camera angle is carefully chosen and any spoken audio must be clearly captured (esp. in the case of language learning videos.)
 Again, viewer acquires procedural knowledge, but it is via vicarious doing and not via showing. In contrast to how-to video, learning videos are only intended to be watched once.

Interview/review: I want to get someone’s view or opinion or tell my own. Can be planned or relatively unscripted. Contain opinions or attitudes. Factual statements are made to support opinions. The visuals may be convenience visuals, but may be planned convey the feel of that person.


Report: Following a script, I report a certain even that happened. The statements I make are factual. The visuals provide depictions of the objects mentioned (broadcast news). In the end the viewer has acquired declarative knowledge.


Documentary: Reports a sequence of events subordinate to some sort of ordering. Temporal ordering is common, or they may be ordered in order to support a thesis. Documentaries include a narrative line: they open questions and resolve them. In the end the viewer has acquired declarative knowledge.


Film (or TV Series): Narrative created for the purposes of entertainment, but may have other elements (didactic, community memory). The basis of film is a complex “contract” between the filmmaker and the viewer, which rests on a series of established conventions that have been developed over the history of filmmaking (the literature traces this system of conventions back into novels). The nature of this contract varies from film genre to film genre. Narratives are created by setting up viewer expectations and then either fulfilling the expectations or failing to fulfill them. Scenes are carefully composed to carry out scene setting, introduce characters and to depict events. Organization can be temporal or otherwise. Shots are set up so that the viewers understand what is going on (in all but a few exceptions, the main action in the shot will be readily visible, for example, the moment that he passes her the gun, the moment that she kisses him). In addition to understanding the plot, the film aims to create a mood for the viewer. Conventions are used to create mood include music, timing shifts (quick shot sequence used to portray the passage or time), lighting, camera angles. Further, films are created to delight the viewers with their film craft, which can involve strict adherence to filmmaking principles (including references to other films) or creative breaks with convention.


Art: With art the contract between the viewer and the creator is not as complex as with film. In fact, it can be considered utterly simple. The viewer simply has to agree, “this is art”. The impression the video has on the viewer is decoupled from the intention of the creator to a greater extent than in film. Art closes the circle and resembles captured video in that it the video is an object in of itself. It assumes a purpose in the act of viewing. Art events are not necessarily depicted so that the viewer understands the “plot”, the conveyance of a certain mood may be highly viewer dependent and there is no narrative. If there is a speech track, there is no predictable coupling between the speech track and the visual channel.



Object: Sometimes we make video and we don't know why. Neither videographer or viewer would readily commit to the "art" label. Even a video that we have made ourselves become objects that inhabits the world of objects, things we come across, think about, try to fit in to the larger pictures. Some of these videos are the ones whose existence in an of itself explains and expands the role of video. Here perhaps there is only a person and a camera and it is not appropriate to speak of intent. Video happens.

In sum, the intent of the creator can be used as a basis of a typology containing many different kinds of video. Each is different from the other in several important respects, including, (1) what information is packaged by the creator into the visual channel and (2) what the relationship is between the visual channel and the audio track. In light of this typology, it is rather curious that we consider multimedia information retrieval to be a single discipline. Instead, every genre presents us with a unique set of challenges -- an entirely different range of issues that need to be face to provide retrieval algorithms that succeed in meeting users' needs.

Thursday, May 6, 2010

Knowing where to search

At the end of last year, VideoCLEF became MediaEval. I thought it was a great name for a multimedia retrieval benchmark evaluation and a mashed up a new logo in an enthused rush. When I needed to go back and find the original illumated "M" that I used, it seemed to be the perfect job for content based image search. I recalled the Best Paper from ACM Multimedia 2009 onVisual Query Suggestion and headed off to Bing image search to try my hand at some combined text and image search.

I quickly found myself wishing that I had more options. In particular, I wanted to chose more than a single image at a time that was related to my query. The VIPER group at the University of Geneva has a Cross-Model search engine that lets you select multiple relavant images for each feedback iteration. You can also select a set of images for negative feedback, which would have been helpful.

But for this particular search, the Bing option to limiting search to black & white proved helpful. After a few iterations, I came up with some nice looking results that gave me a sense that I was really moving the right direction.

However, my search did not return the "M" that I had originally used. I went to Google images, formulated and reformulated. "Letter M illuminated", "Medieval manuscript M", "Illuminated medieval letter"...nothing seemed to help. Arg! Isn't this task easy? Shouldn't this just be duplicate detection?

Then I remembered that when I was looking for the original "M" I wanted to make sure that there would be no licensing issues so that MediaEval could use it freely. I had been experimenting at the time with the Creative Commons search engine so I went back there and put in the simplest of all possible queries "Illuminated M."

Bingo. The original M from the Chronica Polonorum on Wikimedia commons.

How often when we are searching do we remember that Web search is all about recall? Multimodal relevance feedback may expand our queries, but it also limits our results. If I weren't engaging in known-item search I would have never known the "M" I was missing. Similarity along radically simplistic visual dimensions is useful, but enevitably something will fall between the cracks. Thankfully it seldom seems to matter, but we shouldn't let our awareness that we might be missing something slip from our consciousness.

The more interesting observation was that the key to re-finding my image was reconstructing the way that I found it in the first place. Not only knowledge about the "M", but also detailed knowledge of where and how I should be looking for it turned out to be critical.

The search process is entertaining in and of itself. I am not going to reveal how much time I was willing to devote to finding that "M" and browsing through the images that Bing came up with as similar. The visual feedback did turn up a useful by-product -- not the direct target of my search: a beautiful high-resolution "M" that should satisfy gripes about the low quality of our MediaEval logo.