Saturday, June 11, 2011

Search and Spirituality

Today, I discovered an interesting segment of a video clip illustrating someone connecting search and spirituality. Search in a broader sense (beyond "information retrieval") does seem to have a lot to do with our belief systems and our relationship to a sense of higher purpose in life. Coming across a tangible example of the connection between finding information and someone's inner spiritual world stopped to make me reflect. I was struck by the implications for the design of user experience with search engines. What responsibilities do we have as scientists in designing our algorithms and our applications if these then get incorporated into the personal, internal process of individual human beings to find meaning in their own lives by connecting with universal truth?

A
t the moment I am doing the final spot check on the development set for the MediaEval 2011 Genre Tagging release. I was checking out a video with the genre label personal_or_auto-biographical, one of the 26 categories that we are using this year.

I started playing this video to get an idea of what it exactly was about and I was amazed to listen to this guy and watch him speaking. Perhaps the reaction dates me. There is just a striking immediacy to it that I was not expecting. Apparently, he's alone in his car, and talking only for himself and for the camera.



To
really not know who this is, or what happened to him later in 2009 when he stopped publishing episodes is a bit of a science-fiction feeling for me. Watching his video, I am caught up in the present moment of someone who I don't know, over two years after that moment actually occurred. This effect is quite contrary to what he himself is describing. He talks about remaining with himself (someone he nearly by definition must know well) in the present moment.

Or is my witnessing of this nameless present-moment occurring in the past actually simply a new kind of being present?

It certainly seems like it exists on some other plane. Although, I jump immediately to considering what it would take to track the guy down. Gerald Friedland gave a talk at our lab last week about Cybercasing,
using geo-tagged information available online to mount real-world attacks. It's fresh in my mind, the array of possibilities for finding someone by following the trail they leave uploading multimedia to the Internet. One video doesn't seem to hurt, but we quickly loose the intuitions for how our uploading behavior might scale -- allowing people to find us on the basis of who we are and in terms of how we are vulnerable.

On the other hand, this yearning to be present in the moment is so universal, so common to so many, that it really doesn't make this guy so special. He's special, perhaps, in that he can operate a camera and get his video online. Also, clearly he has the gift to generate a speech stream that other people then identify as reflecting their own inner processes. But he's specialness ends in a certain way right there. What he is saying in a way so intensely personal that it once again becomes universal -- it's simply what we look like on the inside -- like the pictures that they show us in grade school of the chambers of our hearts and the insides of our large intestines. This video was in that sense made to be lost in the multimedia avalanche of the Internet.

The guy mentions a name in his metadata, Eckhart Tolle, and I followed the trail and very quickly realizing, by clicking into an Eckhart Tolle YouTube video, that Eckhart Tolle is who my mystery guy is talking about rather than who he is himself. That brought a smile, since this distinction is one that we've previously observed as important for speech media [1].

I listened to Eckhart Tolle for a bit, pondering the metaphor involving the universal similarity of people's large intestines. All of a sudden
Eckhart Tolle is saying, "The mind even started to look at ads for flying back to England, fares, and then the impulse came..." He's sort of hesitating, so you wonder if he's also finding this a little strange, but for me it just seemed like a moment that search for information is playing a clearly in central role in what we would otherwise call our own internal states that make up part of our spirituality. It's the kind of search that we would do nowadays with a search engine.

Eckhart Tolle goes on to talk about "obedience to what came out of the present moment"...it guided his decision making process on where to be when. He goes on to say, "...don't do it on an impulse that is a restless impulse or comes out of any kind of negative emotion". If people listen to what he is saying, and a lot do, and if they combine their search for information with interaction with search engines, I land at the following conclusion: our individual spiritual development is not disconnected from our search engines and especially not from our experience of interacting with them.

In the end, the reason I blog about this might just be that I want to use the YouTube link to that Eckhart Tolle video that will take you right to the jump-in point that I am writing about
http://youtu.be/K1_R3uKJOB4?t=4m18s Goodness knows how much time I've spent discussing video fragment linking and trying to get research money to work on it as a searching speech problem -- I really get a kick out of being able to link into the stream.

We are late releasing the data for the MediaEval 2011 Genre Tagging task. The initial delay was small, but then other things just got in the way compounding the situation. Today, I am trying to be very present in the moment, in order to ignore the stress that I feel about being so late and be very careful about getting the release right the first time around.


And today's experience reminds me of how careful we need to be in all our research. If our search engines are part of our spiritual worlds, we need to design our algorithms and applications with awareness of their potential impact on the trails that we following in our paths of personal development and on the collective, common digestive system of humanity.

[1] Besser, J., Larson, M., Hofmann, K., Podcast Search: User Goals and Retrieval Technologies, Online Information Review: The international journal of digital information research and use, Vol. 34, No. 3, pp. 395-419, 2010.

Saturday, June 4, 2011

LikeLines: Crowdsourced intelligent mulitmedia player

Today we're doing the putting the final touches on LikeLines: Video highlights via web-scale aggregation of moments that viewers like, our entry to the mozilla Drumbeat Unlocking Video challenge, which is closing tomorrow.

The challenge addresses the question "How can new web video tools transform news storytelling?" Our answer to this question is the paradigm of distributed directing that allows news reports to be generated automatically, but without a central reporter. The raw material is footage captured by individuals with cameras and mobile phones who witness an event. One challenge faced is how to filter this footage: in particular, how to find the most interesting points? LikeLines gives the answer to this question.

The LikeLines concept is basically a heatmap that shows how many people found certain portions of a video interesting. You use it if you don't want to watch a video all the way through. Instead, you click the heatmap to jump in to just to the places that are worth watching start watching from there. What's worth watching is decided on the basis of what other viewers found worth watching -- either they tag those segments explicitly by clicking a "like" button or else they let the player record their stop, starting and cuing behavior. We also want the player to be able to make use of multimedia content analysis (visual analysis or speech recognition) in order to be able to "seed" interesting moments. This sort of seeding user contributions with multimedia content analysis has been used by our colleagues:

Ewine Smits and Alan Hanjalic. 2010. A System Concept for Socially Enriched Access to Soccer Video Collections. IEEE MultiMedia 17, 4 (October 2010), 26-35.

Our entry is in the form of video:



The video has been finished for a while now, and now I am just adding some text to make it clear that the idea is elegant, but also quite clever in that you combine user input and multimedia content analysis, which allows you to bootstrap from raw video.

It's a little crazy trying to write, because I need to switch out of research paper mode in to the mode of "hey, this will really work" and "hey look everyone, this is totally needed, totally non-trivial and totally does not exist anywhere yet". I am working now (when I stopped to write a blog post) on a sentence communicating that we can address the cold start issue with content analysis based seeding. And that verification using content analysis will help to control spam. And that if all goes well the whole thing should be able to learn by itself: It will require some R&D effort, but all the pieces of technology needed already exist.

Also, I had a little bit of trouble getting the right tone for the biography. So we're big shots at a cool technical university in the Netherlands? I guess that's important to communicate. But how to say that we are also passionate about supporting distributed and democratic news? Do I divulge that the first draft of LikeLines was churned out on a bus from Boston to Portland, the video was recorded in a long after-hours effort, and the whole thing has been discussed in every detail in chat sessions?

And how to communicate that we are doing this because it's what we love to do? We had some light-hearted lines in the bio to convey this tone (about our cat-video habit on YouTube and about me largely eschewing social media for the traditional postcard), but those got dumped in favor of some harder hitting facts about our experience in this area: right people, right skills, right place, right time...

In the end, I'm also in this to experience the crowdsourcing aspect of working on innovation in a open collaboration environment. What a breath of fresh air in the daily grind of publish and perish. And the giddy joy of communicating a concept that is ripe, feasible and useful.

Sunday, May 29, 2011

Tagging Love and Affection: Part II

Perhaps even the more interesting thing about the wedding in terms of modern media was the interaction between the professional photographers and the wedding guests who were taking pictures. It seemed like these were two completely different activities in terms of the results that they were aiming to produce. The photographers created an amazing album of storybook moments -- and the guests -- well, speaking for myself at least -- took pictures of people as people that they knew.

The professional photos actually included shots of members of the bridal party and guests photographing the bride and groom and each other. There was a particular dramatic one of the best man from the back taking a picture of the groom and you can see the groom twice: once over the shoulder of the best man in the display of his cell phone and once sitting as the main subject of the image.

There is another one in which one of the bridesmaids is taking a picture of the newly weds. It's like this act of greeting, an 'I was there with you in your moment of bliss and I was so so happy for you.' Taking a picture is like smiling, waving or winking at someone -- except that it is asynchronous, delayed in time. From the past, a shout out, "Congratulations!"

I was struck how the professional photographers were able to use the act of photo-taking as a way to depict the love and affection between friends and family members. Its not just the places that we tag with our photos, as I've discussed in a previous blog post, but its people, too.

And this is where the difference between the professional photographers and the guests really became apparent. I was using the little camera of my mom, so my pictures of the wedding are not qualitatively speaking very good. It would probably take some improvement of my photographic skills and not just a better camera to get high quality pictures.

But there was something that struck us about them. I naturally looked for the people in our family who we see in frequently and took pictures of people talking that only get to see each other once every several years -- if at all. I took pictures of people holding the youngest baby in our extended family -- that capture the moment that generations within the extended family meet for the first time. I took pictures that showed siblings engrossed in conversations with each other -- showing how the intensity of how we speak and how we listen. The photographers didn't know us and although the wedding pictures were beautiful, aesthetically not to be surpassed, we really love to look at the personal pictures, because they are somehow more "us".

Maybe the "real" pictures are the pictures that we take that mark love and affection. They encode our personalities, our common past and our hope for the future.

The implications are quite large for the field of multimedia retrieval, as revealed by the following line of reasoning: Life is finite. We only live so long and can support close relationships with so many people. If Dunbar is right it is a very limited number indeed. If we take photos at particular moments, such as moments of expression of affection that I am describing here, then the number of total pictures that we take is also limited. If we keep on insisting that our multimedia retrieval algorithms must be able to handle millions and millions of photos, then we run the danger of missing out on developing important techniques. Multimedia algorithms developed for relatively small numbers of photos can afford to be computationally more complex. If we ignore the "small set" problem, we run the danger of not developing the best possible algorithms to personal multimedia retrieval challenges.

As final comment, I can add that during the wedding I was already challenged by an image retrieval problem. I took maybe 200 photos. I wanted to show a special photo of my mom -- taken a few minutes back -- to the cousins I was sitting with at the dinner table. It took me so long to flip through the index to find those photos on that small screen. Very disruptive for dinner conversation.

I had the idea that I was the one that should have been wearing the GSR sensor. I am sure that my affective peak was physiologically measurable when I took the picture of my mom, saw it on my camera display for the first time and realized I had gotten a once-in-a-lifetime shot. If my camera display could take me right to peak pictures, it could be a much more functional device: transcending a capture to also support storytelling as well.

On second thought skip the GSR. I'm sure I jumped up and down. The accelerometer on a mobile phone could have picked that up. If all else fails, make a photo taking app that encourages me to shake the thing when I notice I like a picture. Better stop blogging and start implementing.

Tagging Love and Affection: Part I

Oh, my baby cousin is all grown up and got married. The wedding was a fairy tale: the kind that they show in the last scene of a good movie where you then sit through the entire credits in hopes that your eyes are reasonably dried up by the time you walk out into the bright public space beyond the theater.

The picture shows the bride's foot. Maybe one would expect a glass slipper, but this is a galvanic skin response sensor that recorded throughout the ceremony so that the bride and groom can later transcribe a mutual emotional trajectory. The sensor records skin conductance, which is related to moisture, and the signal must have been off the charts -- it was a wonderful wedding and but it was also a beautiful sunny day. Sunny in the intense sense, where you gain an practical understanding of why the British royal family feels that weddings require large-brimmed hats.

The nuptial couple must have been perspiring at least lightly and I suspect their respective sensors registered one continuous high throughout the proceedings. And indeed: I wish them this for their married lives that their joint signal continuously reflects a wonderful life experience. Or, if it's not a continuous high, may they at least be gifted with the ability to recognize any large dips as outliers and ignore them.

Lasting love is a complex phenomenon and I suspect that the most reliable expressions of it are more under our conscious control. In the next post I relate it to ... picture taking!

Saturday, April 23, 2011

Proto-semantics: Where cognition meets language

A recent Nature paper entitled Evolved structure of language shows lineage-specific trends in word-order universals has made quite a splash in the media and has some important implications for the study of linguistics and application of linguistic principles. The splashy aspect of the paper is that it allows the media to declare the discovery of purported fundamental flaws in the theory of Universal Grammar.

The more subtle implications of the research are highlighted in an Wired Science article entitled Evolution of Language Takes Unexpected Turn. This article quotes Michael Dunn, author of the Nature paper, as stating, “What languages have in common is to be found at a much deeper level. They must emerge from more-general cognitive capacities.”

This finding suggests that if we indeed are interested in understanding our human language capacity, it is time to return to again give serious consideration to proto-semantics, the pre-lexical dimensions of meaning arising from the interaction of our human perception and the real world in which we live. The commonality underlying the way we speak arises from the similarity of our cognitive processes, which in turn can be traced to the structure and function of our brains. The way in which proto-semantics is encoded into the wide range of human languages and the variety of the evidence that we can observe of its existence are re-emerging as the more critical research questions. Whether we need a universally applicable syntactic mechanism to explain language may at this point be less relevant.

As a scientist, I personally renegotiated my relationship with Universal Grammar 15 years ago while doing the field work in theoretical linguistics on Baule in Cote d'Ivoire. No amount of pushing and shoving would fit the verb series constructions that I was studying into a framework determined exclusively by settings of universal language parameters previously identified in other languages. It appeared to be a classic case of West Africa Wins Again. My laptop held up wonderfully under humidity, sand and travel in the Cote d'Ivoire -- but my fine theory failed miserably.

One aspect of Baule that I looked at in detail is the conditions under which it is possible to omit a pronoun object in a sentence. In answer to the question "Did you beat the drum?" one replies "Yes, I beat", which must be translated into English as "Yes, I beat it". You leave out the pronoun because (simplifying somewhat) the drum is an object whose purpose is to be beaten (in the sense of played). The functional relationship in the real world as perceived by humans makes the verb and its object inextricable with respect to their meaning. The pronoun is considered redundant and can (in fact, must be since the effect is encoded into the grammar of the language) left out.

Contrast that with "Did you see the drum?" and the answer "Yes, I saw it". Here, the acting of seeing and the drum itself do not have a tight connection in the cognitive conceptualization of the world. The drum is seen because it has a physical existence, but it is not explicitly designed to be seen nor can it be considered an object typical of the sorts of objects that are seen by humans. In the answer "Yes, I saw it." it is not possible to omit the pronoun.

The syntax of Baule failed to fit into my neat universal framework, but I learned an extraordinary amount about the range of phenomena that "register on the radar" of human cognition and naturally also find their way into language. Other languages express these phenomena in other ways, or chose not to express them at all. However, that proto-semantic generalizations can be found to exist at all and which generalizations exist is both interesting and useful.

What does this have to do with search (the topic of this blog)? If proto-semantics is a reflex of human perception of the real world, we should expect that its influence impacts not only relatively low level syntactic processes, but also the more complex ways in which humans communicate and exchange information: information creation, curation and retrieval. If we can push forward this line of inquiry, a full answer to this question will surely be important enough to make its own splash.

Sunday, April 17, 2011

From the Rialto: On the Relevance of Geotags

This weekend I am in Venice in between an Networked Media and Search Systems Concertation meeting and ICMR 2011. I am in the processing of preparing the slides for a talk [1] that I will give on Tuesday morning on tagging and geotagging. Once again, I find myself busy with the question of what does it mean for a geotag to be relevant to a photo.

At first blush, the answer to this question is obvious: A geotag is relevant to a photo if it gives the exact longitude and latitude of the position of the camera when the photo was taken. I have this information for the photo above, which was taken earlier today -- but it reflects so little of the larger story. On the one hand our cameras (in this case we're using a new Android phone) pin us down and on the other they tell not much at all.

I originally asked this question about the relevance of geotags after the talk of Pavel Serdyukov at SIGIR 2009, "Placing flickr photos on a map." I garbled the formulation, Pavel didn't understand it, and I have since come to understand the large number of dimensions involved in the answer. "Where does a photo truly belong on a map?" isn't an obvious question and doesn't have an obvious answer, either.

It's generally accepted that knowing the position of the photographer underspecifies the location of the photo since it is not known in which direction the camera was pointing or on what it was focused. It could possibly be the case that the subject of the picture is quite far away from the photographer, making the geo-coordinates of the photographer not the most relevant position to the photo in terms of the depicted content.

The photo from today is a good example. It was taken in Venice, but it is also relevant to Trento, since I am here in Italy on a trip to Trento and only happened to come to Venice for the weekend. I will remember the trip as a "Trento" trip. But the Trento association goes beyond my own personal sense of the proper place of the photo -- anyone who was on a business trip to Trento might find Venice related since they are possibly flying into a Venice airport or like myself also planning a touristic stay in Venice.

The sign in the background says "Per Rialto" and is pointing the way towards the Rialto bridge. We're on our way to the bridge and have taken the photo along the way -- it is relevant to the geocoordinates of the Rialto bridge in the sense that this was the final destination of the walk.

The photo was taken for my family back home. My great uncle is sick and in the hospital at this moment. He is a central node in my extended family and ties us together across space and also across time -- and for years has passed along to us the lore of the generations that preceded us -- a mixture of stories, poetry and family values -- a nourishing flow that sustains us and binds us together. One of the things that he does is write email newsletters to the family members and this newsletter has a section that is called "From the Rialto", which describes the current activities of various family members. Here I am sending my own thoughts out from the Rialto to my family back home and my great uncle in the hospital. The picture has more to do with them and where they are than here.

Further, I don't really remember why the section of the newsletter is called "From the Rialto". It might not be the Rialto in Venice (which is actually an entire district and not just a bridge). It might be an entirely different Rialto. My picture is actually related to the original Rialto used by my uncle, except that I don't know where that is. I will admit that there are probable not a lot of photos out there that enjoy this sort of relationship to place -- but the main point is the existence of possibly relevant places that the person who takes the photo is not necessarily aware.

Actually, I just wanted a picture with a Rialto sign in the background. These signs are numerous in Venice...in fact they sell T-shirts on which they are depicted. In a certain way, the photo is a "Woman in from of Rialto sign" picture and is also relevant to anywhere else a Rialto sign is located. In a way, I am standing at place that has many different geo-coordinates.

The renown of the Rialto bridge is attested by the numerous replicas that exist, for example, in Las Vegas. The original and its imitators are related -- photos of one are of interest to people visiting the location of the other. It would be useful for a picture of the Las Vegas Rialto bridge to bear a tag that relates it to the one here in Venice.

I briefly entertained the idea that it is generally accepted that the relevant geocoordinates for a photo are the ones at which the photographer is standing, since those geocoordinates would be considered relevant to the widest range of people, i.e., be the most objective. Indeed, this is the relevance criteria for geotags that we apply within the MediaEval Placing Task. My association of the photo with the folks back home is relevant to a handful of souls, whereas anyone could recognize the connection with this corner of Venice. However, replicas are part of a common shared culture across the world and the connection between an image of a replica and the geo-coordinates of the original could be considered of possible interest to anyone and not just a personal circle.

Finally, there is the question of the relevant geographical resolution. For me, this is a "Venice" picture. I don't need to know the exact corner again -- nor will I try to find it. We overheard a group of Americans talking about their vacation -- and noticed that they referred to themselves not as being in Venice or in Italy, but in Europe. Europeans hardly ever reflect upon the fact that they are in Europe...they take it for granted.

Wandering through Venice its tempting to just fall into people watching -- making up small stories about couples out enjoying the good weather and the beautiful city: how they met, how they fell in love and what the future has in store for their mutual happiness. The relevance of the geo-coordinates of the photographer is obvious for those who want to take people watching to the next level -- if the couple standing next to us a bridge near San Marco uploads their pictures to Flickr, we will know even then more about them...we could check if their trip to Venice provided the solid basis for a stable relationship, or whether they ended up parting ways. Taking the people watching to the extreme: It would finally be possible to check whether kissing under the Bridge of Sighs does its work in guaranteeing everlasting bliss.

[1] Larson, M., Soleymani, M., Serdyukov, P., Rudinac, S., Wartena, C., Murdock, V., Friedland, G., Ordelman, R. and Jones, G.J.F. Automatic Tagging and Geotagging in Video Collections and Communities. ACM International Conference on Multimedia Retrieval 2011.

Sunday, March 27, 2011

Your mistrokes are costing me keystrokes

My mother insisted that I learn how to touch-type in high school and so I spent a summer after ninth grade or so pecking out the obligatory practice sequences.

At times I've caught myself being vain about my ability to take notes in lecture on my laptop and still make eye contact with the speaker. That's balanced by the many meetings in which I've taken minutes because it was uncomfortable to watch my colleagues engaged in seek and hit missions to input even simple sentences. And also by colleagues who are utterly unfazed by my scripting proficiency, but who light up when I am able to turn to smile as they walk in without breaking my stride in the e-mail I'm writing, "Hey, you can look at me and still type!"

Well, Google has found a way to keep me humble and to break down any remaining feelings of typing elitism. I type in a perfectly formed query, work meaning wired and Google finds it necessary to give me results for word meaning weird (check out the screen shot). My query is meant to find me the Wired article that I read in the UK edition on Thursday about a study that had been done at Duke about the meaning of work. I want to have another look at it so see if I can relate it to issues of worker engagement in crowdsourcing -- work that I currently collaborating with my colleagues on.

Evidently is what is happening is that the words "work" and "wired" are mistyped so frequently by searchers, especially in the context of the word "meaning", that the probably of me actually having meant to type work meaning wired is vanishingly small, or at least has fallen off the bottom of the Google charts.

I permit myself a half formulated cuss, well under my breath, and then I type in "work" meaning "wired".

OK. Well, my complaint about the wasted keystrokes may not be my only complaint: the extra apostrophes don't get me the article I want either.

But I do serendipitously stumble across a blog post urging scientist to move beyond papers and discuss the importance of their findings in blogposts: http://www.wired.com/wiredscience/2010/09/does-all-scientific-work-deserve-public-attention/ Hmm. So maybe I should be saying actually more why I want the paper about feelings of meaning at work and how it relates to what we are currently doing in the area of crowdsourcing. Our latest crowdsourcing paper:

Vliegendhart, R., Larson, M., Kofler, C., Eickhoff, C. and Pouwelse, J. Investigating Factors Influencing Crowdsourcing Tasks with High Imaginative Load Proceedings of CSDM 2011 (WSDM Workshop on Crowdsourcing for Search and Data Mining), 2011.

reports on a study whose results suggest that work quality is improved when workers are asked to provide explanations for the answers that they provide. In particular, they remain more engaged if they are required to provide an explanation for a personal preference that they have expressed. I would like to try to relate this to worker perceptions of meaning of the work.

In the end, I get up from behind the computer and walk the five meters to the chair where the hard copy of UK Wired was waiting and flipped quickly to the page, "At Duke University, when we have run experiments on feelings of "meaning" at work (tinyurl.com/4rwkg85)..." there it is.

One last go at Google, I find the article online by typing in the question asked in its final sentence, "How can we enhance the feeling of meaning in those around us?"

Maybe I'll start by being more patient with sloppy typists.