Sunday, November 22, 2009

Lemonface

Since I was going to be mentioning Flickr Video in my presentation about VideoCLEF at last week's TRECVid 2009 workshop, I decided I should try and make one. This video was shot out the window in the morning. And yes, I dutifully made sure it was properly geo-tagged as Gaithersburg. Sure indeed, capturing the ripple makes the flag image come alive..."It's like a photo, but it moves!"

Is Flickr video a long photo?

Maybe it's even a bit too alive. The ripples would be more dramatic had the camera been stabilized...you can see I'm not holding it still. Where does it end? Interactive moving pictures, a la Harry Potter's newspaper, I suppose.

Immediately after downloading this video off of my camera, I lost it. My first reaction: Had I downloaded all of the pictures on the camera? (Yes!) Only then did I start fretting about how much money I paid for the camera, a Cannon PowerShot SD1100, and how much it would cost to replace it.

I felt it viscerally: it's not just a bunch of bits! Content matters to us in a very personal way. Since I didn't empty the camera memory card, are my photos are floating around out there, in whatever space that the photos on lost cameras go. What if this space turns out to be the Internet? Would I be embarrassed?

In my opinion the most embarrassing picture on the camera is a lemon-face from Jay and Silent Bob. And yes, there is also a corresponding lion-face out there...somewhere. I have this most pressing information need: "Where is my camera?"

Friday, October 23, 2009

SSCS 2009


The Third Workshop on Searching Spontaneous Conversational Speech (SSCS 2009) took place on 23 October 2009 in Beijing China in conjunction with ACM MultiMedia 2009. We had a great set of demos and talks. As an organizer this gives you a warm pleased feeling -- all that work is worth it. Domains covered included broadcast, meetings, interviews, telephone conversations, podcasts and voice tagging for photos. The approaches presented involved using a variety of techniques including subword units, exploiting dialogue structure, fusing retrieval models, modeling topics and integrating visual features. Such events serve to highlight the importance of the spoken word in many multimedia access and retrieval applications. And also to remind us how far we are from exploiting it fully.

Thursday, October 22, 2009

Well, behind every joke there's some truth

The winner of the ACM MultiMedia Grand Challenge at ACM MultiMedia 2009 was "Joke-O-Mat: Browsing Sitcoms Punchline by Punchline." This application uses speaker diarization and laughter detection to annotate sitcoms and present them to the viewer in an interface that allows presents jokes ranked by laugh reaction, grouped by character and associated with context. Joke-O-Mat underlines the importance of the speech track for multimedia access.



What to do when multimedia doesn't contain a laugh track? In VideoCLEF 2009 we ran a narrative peak detection task. The goal was to detect points in videos where viewers perceive heightened dramatic tension. Today, CLEF working notes, tomorrow our own Peak-O-Mat?

Thursday, September 10, 2009

Getting the words right

The final day of Interspeech 2009 here in Brighton. It's been a great conference and each and every keynote has been well worth getting up for. This morning, Mari Ostendorf talked about "Transcribing Speech for Spoken Language Processing." Interspeech encompasses a staggeringly broad spectrum of perspectives on speech research and technology. For every point here, there is an immediate counterpoint, and it was without doubt under influence of this chorus that the opening slide of the keynote this morning displayed a long-play version of the title reminding the audience that they would be hearing about transcribing human-directed human speech, as opposed to speech that humans produce to communicate with computers.

The message from the keynote that will ring longest in my ears was, "The goal of speech transcription is information access." This leaves open of the course, the question of what is the information and what is the access when it comes to content that contains the spoken word. I find myself compiling little lists of domains in which information encoded in spoken audio could be important: podcasts, video diaries, lifelogs, meetings, call center recordings, social video networks, Web TV, conversational broadcast, lectures, discussions, debates, interviews and cultural heritage archives, home videos, photo annotations, video conferences. These lists invariable end with etc. etc. etc. And what constitutes access (keyword search, retrieval, question answering, browsing, recommendation...) is another question to which we can't give a closed-set answer.

My personal experience doesn't really support the idea that we need to push the envelope. The last video I watched I found because a link was sent to me by my cousin. The content of the video was a short clip of her new cat purring. No real access problem there. No information either. The purr did not inform me in the conventional sense. In fact, there wasn't much human speech involved at all. Nonetheless, I found the content supremely worthwhile of my watching time. Although my own multimedia access needs are a string of examples of this pre-solved sort, I do agree that the challenge of access to speech-based information is a serious one and will require a great deal of effort to address.

The full phrase in Ostendorf's slide read, "The goal of speech transcription is information access, not just getting the words right." But maybe it is about "getting the words right". The words referred to are, presumably, white-space delineated grapheme strings, lexical words, citation forms. But we can also see a word as the totality of knowledge that a human needs to possess in order to deploy it in human-to-human communication. There may be a limit to how far we can go beyond that sort of word and still remain within what is meaningful in the context of our information access needs.

We can go for prosody, for speech act, subjectivity, affect, but in the end we'll never capture the "you had to have been there" component of understanding. And already the moment of that particular purr video has passed and my next need for video content will be for a new one.

Friday, August 7, 2009

Say Anything


Standing drinking a Diet Coke I gazed at an advertisement for the Guardian outside the Spar on the campus of Dublin City University.

The advertisement stated, "Owned by no one, free to say anything." I paused. My paper of choice is associated with the motto "All the news that's fit to print." Never worried about it before, but in comparison, it suddenly seemed a bit dated.

It's relatively uncontroversial to consider source when assessing the credibility of media. In our work on the PodCred Framework we cite Rubin and Liddy (2006) as a source for the notion that user generated media builds credibility by avoiding hidden bias. I smiled at the idea of The Guardian as a huge blog; and then again at myself for finding that funny.

The PodCred Framework includes an indicator meant to capture the source of the income of the podcaster: stores, sponsors, advertisers. The idea that transparency of funding does indeed impact listener satisfaction with podcasts hasn't been test driven yet, too my knowledge. But seeing the Guardian sign made my thoughts return to consideration of its potential.

Then I finished my break and went back inside to continue working on VideoCLEF assessment management tasks, which is why I am currently in Dublin, and my mind turned to other things.

I've found the image to accompany this post at http://www.scaryideas.com/ If this is indeed a scary idea, I wonder if it indeed sells papers. But if it's really owned by no one, perhaps they need to make the link up to who is actually doing the writing. (Note to self: I do, too.)

Friday, July 24, 2009

Preferential Attachment at SIGIR 2009

Waiting for my plane yesterday at Boston Logan airport I found myself starting out the window and reflecting on scale-free networks rather than attending to my e-mail as I should have been. The reflections were, of course, inspired by the keynote of Albert László Barabási entitled "From Networks to Human Behavior" and consisted mostly of wondering why the properties of scale-free structure are perceived as unintuitive or unnatural. A conspiracy is improbable, but can it be that we are subtly taught to expect behaviors characteristic of random connections? Barabási used the word "democratic" in describing random networks and one of the audience questions afterwards dealt with how to overcome potential isolation between nodes in scale free configurations. Yes, they are out there, but how do we fix them? The answer was, to paraphrase, know it's there and work with it. I'd been standing in line with an audience question of my own, but I resumed my seat, deciding that it was basically covering the same ground: "Are the scale-free networks in and around us inherently irreconcilable with our notions of democratic forms of organization?"

Later that evening, in a semi-conscious effort not to always hang with exactly the same crowd, I fell in with group of bloggers for an interval that would, unknown to me as it was happening, later be referred to as Day 2 post banquet. Face-to-face discussion with bloggers seems to have helped to counter my conviction that I am almost, but not quite, entirely unlike a blogger myself. Except for my apparently innate repugnance for preferential attachment. Work with it.

Monday, June 1, 2009

The Netbook Effect

"Give a laptop. Change the world." are the welcoming words of the website of One Laptop Per Child. A Wired article entitled "The Netbook Effect: How Cheap Little Laptops Hit the Big Time" (Wired, March 2009) tells the story of how the low cost no-frills netbook evolved from the One Laptop Per Child initiative and predicts that netbooks will constitute 12% of the laptop market in 2010.
It sort of takes my breath away that I might actually live to see One Laptop Per Child happen. It gives me hope that we might eventually move not only as individuals, communities and countries, but as an entire species beyond subsistence level concerns.
At the same time, another part of my brain is formulating questions about what happens as an increasing proportion of personal machines are "thin" in the sense that they have low processing power and low storage capacity. The vision of democratizing the supply of services with high computational loads by distributing them over private citizens with unused processing capacity may have to be abandoned. Do we want to encourage a future in which we must rely on a distant center to store and swap our video content? Do we want to close the door on the possibility of internet content search that is supported by a million modest contributions of storage space and cpu cycles? Are we ready to give up our private capacities and allow resource rich hands to further accumulate a power monopoly removed beyond influence of the individual?