Thursday, July 28, 2011

It's complicated.

I just logged into the manuscript management system of a journal that will remained unnamed and was greeted by this information icon plus message.


Is this system really user friendly or is it exactly the opposite? The page goes on to state, "So please do look at the .... hints and warnings that we’ve put in the green box at the top of each page. They will save you time and ensure that your work is processed correctly."

I identify with this sentence. I am always telling people that they are probably not going to understand everything I am saying. It sounds like standard US-English, but I speak fast, with a lot of unexpected vocabulary, large dynamic range and obscure cultural references and I have just enough of a regional accent. People who easily follow fast paced Hollywood movies think they should understand me. But really, I tell them, the skills just might not transfer and it might not be your fault.

I started warning people after years of trying to slow down and choose simple vocabulary. I still try to do this in formal, large group situations with people I don't know very well. However, there are just too many people that insist on speaking English with me: the other languages I speak are drying up and dying and my English will go that way as well if I don't occasionally take it for a walk and deploy it in its full erratic richness. Somehow, I do really identify with this system: "Hey, World! I'm complicated. Deal with it."

And, like this system, I ensnarl myself in some sort of paradox of self-reference. If I don't speak in a way that will allow people to understand me -- how can I expect them to grasp that I want to communicate to them that I might be difficult to grasp? Will I ever succeed in motivating anyone to bring up the extra dose of patience and attention that it takes to follow me? Won't they be suspicious of me if they know I know that I am difficult to understand and also that I am not doing anything about it? Can I ever convince anyone that the effort is worth the payoff?

Likewise: Can we except that the system is indeed being self-explanatory when it claims of itself that it is not self-explanatory? Is it worth the effort and patience it takes to cut through that knot? In the end it's a matter of trust. I smile at the message and the information icon and consider how to move forward.

In the end, I decide to whip off a blog post, sighing when the word "ensnarl" can't be typed without a red misspelling line under it, and return to my attempt to extract the manuscript I'm supposed to be reviewing from the system.

I now sympathize with the interface. The message has worked: I've agreed that it's ok to be complicated.



Wednesday, July 20, 2011

Google doesn't love me.

As an IR researcher, I tend to obsess about why Google can't always deal with my queries. I fall into this bad habit when I have a lot of better things to do with my time and even when I know it is not getting me anywhere.

Today I needed to recall the details of the relationship between Wikipedia and Wikimedia Commons so that I could get it just right for a text I was writing. I typed "relationship between wikipedia and wikimedia commons" into the Google search box and was rewarded as my top hit the link on the image pictured at the right. The rest of the list was of the same ilk.

Oh, my gosh, why is Google reacting this way? This is not how I expected by evening to play out! Actually, "relationship" was sort of information that I wanted and "Wikipedia" and "Wikimedia Commons" were the two entities whose relationship I wanted to understand. Wasn't I making myself clear?

Is Google interpreting my named entities being as the source of the information? Is Google trying to tell me something? Such as, I should be reading Wikipedia instead of writing about Wikipedia?

Is Google trying to gently point out to me that I should be doing more image search? Or maybe giving me subtle support for my opinion that there is a very fine line between navigational and transactional queries?

Is Google relating my query to religion in order to express support for my blogpost last month on Search and Spirituality, which was written when I was in rather a strange mood? Does Google want to evoke in me again that vein of reflection?

But there is an alternative to this vein of inquiry: a simply, very plausible explanation. It runs closely along the lines of the now infamous: "He's just not that into you". Google doesn't do what I would anticipate or what would make be satisfied and happy because Google simply doesn't love me. The evidence is there: my information needs remain unmet and my search goals unreached.

Google and I obviously have a relationship problem. It goes beyond my searches on "relationships". But yet, I keep on returning to that tempting search box again and again. Do I have some sort of a genetic predilection with maintaining a dysfunctional relationship with my search engine? At least until I find an alternate outlet that does happen to be "into" me and my information needs.

In the meantime, I guess I can go and find myself a copy of the Gospel of Mark: I never realized that there was so much overlap.

Saturday, July 2, 2011

Crowdsourcing Best Practices: Twenty points to consider when designing a HIT

Your very first glance at worker responses on the very first first task you crowdsource tells you that there are very different kinds of workers out there in the crowdsoucing-sphere, for example, on Mechanical Turk. Some of the responses are impressive in the level of dedication and insight that they reflect, others appear to flatly fail the Turing test.

It is also quite striking that there are different kinds of requesters. Turker Nation gives us insight on to the differences between one requester and the next, some better behaved than others.

What is particularly interesting is the differences among requesters who are working in the area of crowdsourcing for information retrieval and related applications. One would maybe expect there to be some homogenity or consensus here. At the moment, however, I am reviewing some papers involving crowdsourcing, and no one seems to be asking themselves the same questions that I ask myself when I design a crowdsourcing task.

It seems worthwhile to get my list of questions out of my head and into a form where other people can have a look at it. These questions are not going to make HIT (Human Intelligence Task) design any easier, but I do strongly feel that asking them should belong to crowdsourcing best practices. And if you do take time to reflect on these aspects, your HIT will in the end be better designed and more effective.
  1. How much agreement do I expect between workers? Is my HIT "mechanical" or is it possible that even co-operative workers will differ in opinion on the correct answer? Do I reassure my workers that I am setting them up for success by signaling to them that I am aware of the subjective component of my HIT and don't have unrealistic expectations that all workers will agree completely?
  2. Is the agreement between workers going to depend on workers' background experience (familiarity with certain topics, regions of the world, modes of thought)? Have I considered setting up a qualification HIT to do recruitment? or Have I signaled to workers what kind of background they need to be successful on the HIT?
  3. Have other people run similar HITs and I have I read their papers to avoid making the same mistakes again?
  4. Did I consider using existing quality control mechanisms, such as Amazon's Mechanical Turk Masters?
  5. Is the layout of my HIT 'polite'? Consider concrete details: Is it obvious that I did my best to minimize non-essential scrolling? But all in all: Does it look like I have ever actually spent time myself as a working on the crowdsourcing platform that I am designing tasks for?
  6. Is the design of my HIT respectful? Experienced workers know that it is necessary for requesters to build in validation mechanisms to filter spurious responses. However, these shoud be well designed so that they are not tedious or insulting for conscientious workers who are highly engaged in the HIT: it is annoying and breaks the flow of work.
  7. Is it obvious to workers why I am running the HIT? Do the answers appear to have a serious, practical application?
  8. Is the title of my HIT interesting, informative and attractive?
  9. Did I consider how fast I need the HIT to run through when making decisions about award levels on also when I will be running the HIT (on the weekend)?
  10. Did I consider what my award level says about my HIT? High award levels can attract treasure seekers. However, award levels that are too low are bad for my reputation as a requester.
  11. Can I make workers feel invested in the larger goal? Have I informed workers that I am a non-profit research institution or otherwise explained (to the extent possible) what I am trying to achieve?
  12. Do I have time to respond to individual worker mails about my HIT? If no, then I should wait until I have time to monitor the HIT before starting it.
  13. Did I consider how the volume of HIT assignments that I am offering will impact the variety of workers that I attract? (low volume HITs attract workers that are less interested in rote tasks)?
  14. Did I give examples that illustrate what kind of answers I expect workers to return for the HIT? Good examples will let workers concerned about their reputations judge in advance if you are likely to reject their work?
  15. Did I inform workers of the conditions under which they could be expected to earn a bonus for the HIT?
  16. Did I make an effort to make the HIT intellectually engaging in order to make it inherently as rewarding as possible to work on?
  17. Did I run a pilot task, especially one that asks workers for their opinions on how well my task is designed?
  18. Did I take a step back and look at my HIT with an eye to how it will enhance my reputation as a requester on the platform? Will it bring back repeat customers (i.e., people who have worked on my HITs before)?
  19. Did I consider the impact of my task on the overall ecosystem of the crowdsourcing platform? If I indiscriminately accept HITs without a responsible validation mechanism, I encourage workers to give spurious responses since they have been reinforced in the strategy of attempting to earn awards with investing a minimum of effort.
  20. Did I consider the implications of my HIT for the overall development of crowdsourcing as an economic activity? Does my HIT support my own ethical position on the role of crowdsourcing (that we as requesters should work towards fair work conditions for workers and that they should ultimately be paid US minimum hourly wage for their work)? It's a complicated issue: http://behind-the-enemy-lines.blogspot.com/2011/05/pay-enough-or-dont-pay-at-all.html
The workers on Mechanical Turk refer to themselves as "turkers". This act of self-naming signals a sense of community, of a common understanding of what they are doing, the commonality of the activity that they are all engaged in.

What do we as requesters call ourselves? Do we have a sense of community, too? Do we enjoy the strength that derives from a shared sense of purpose?

The classical image of Wolfgang von Kempelen's automaton, the original Mechanical Turk, is included above since I think it sheds some light on this issue. Looking at the image we ask ourselves who should be most appropriately designated "turker"? Well, it's not the worker, who is the human in the machine. Rather it is the figure who is dressed as an Ottoman as is operating the machine: If workers consider themselves turkers, then we the requesters must be turkers, too.

The more that we can foster the development of a common understanding of our mission, the more that we can pool our experience to design better HITs, the more effectively we can hope to improve information retrieval by using crowdsourcing.

Saturday, June 11, 2011

Search and Spirituality

Today, I discovered an interesting segment of a video clip illustrating someone connecting search and spirituality. Search in a broader sense (beyond "information retrieval") does seem to have a lot to do with our belief systems and our relationship to a sense of higher purpose in life. Coming across a tangible example of the connection between finding information and someone's inner spiritual world stopped to make me reflect. I was struck by the implications for the design of user experience with search engines. What responsibilities do we have as scientists in designing our algorithms and our applications if these then get incorporated into the personal, internal process of individual human beings to find meaning in their own lives by connecting with universal truth?

A
t the moment I am doing the final spot check on the development set for the MediaEval 2011 Genre Tagging release. I was checking out a video with the genre label personal_or_auto-biographical, one of the 26 categories that we are using this year.

I started playing this video to get an idea of what it exactly was about and I was amazed to listen to this guy and watch him speaking. Perhaps the reaction dates me. There is just a striking immediacy to it that I was not expecting. Apparently, he's alone in his car, and talking only for himself and for the camera.



To
really not know who this is, or what happened to him later in 2009 when he stopped publishing episodes is a bit of a science-fiction feeling for me. Watching his video, I am caught up in the present moment of someone who I don't know, over two years after that moment actually occurred. This effect is quite contrary to what he himself is describing. He talks about remaining with himself (someone he nearly by definition must know well) in the present moment.

Or is my witnessing of this nameless present-moment occurring in the past actually simply a new kind of being present?

It certainly seems like it exists on some other plane. Although, I jump immediately to considering what it would take to track the guy down. Gerald Friedland gave a talk at our lab last week about Cybercasing,
using geo-tagged information available online to mount real-world attacks. It's fresh in my mind, the array of possibilities for finding someone by following the trail they leave uploading multimedia to the Internet. One video doesn't seem to hurt, but we quickly loose the intuitions for how our uploading behavior might scale -- allowing people to find us on the basis of who we are and in terms of how we are vulnerable.

On the other hand, this yearning to be present in the moment is so universal, so common to so many, that it really doesn't make this guy so special. He's special, perhaps, in that he can operate a camera and get his video online. Also, clearly he has the gift to generate a speech stream that other people then identify as reflecting their own inner processes. But he's specialness ends in a certain way right there. What he is saying in a way so intensely personal that it once again becomes universal -- it's simply what we look like on the inside -- like the pictures that they show us in grade school of the chambers of our hearts and the insides of our large intestines. This video was in that sense made to be lost in the multimedia avalanche of the Internet.

The guy mentions a name in his metadata, Eckhart Tolle, and I followed the trail and very quickly realizing, by clicking into an Eckhart Tolle YouTube video, that Eckhart Tolle is who my mystery guy is talking about rather than who he is himself. That brought a smile, since this distinction is one that we've previously observed as important for speech media [1].

I listened to Eckhart Tolle for a bit, pondering the metaphor involving the universal similarity of people's large intestines. All of a sudden
Eckhart Tolle is saying, "The mind even started to look at ads for flying back to England, fares, and then the impulse came..." He's sort of hesitating, so you wonder if he's also finding this a little strange, but for me it just seemed like a moment that search for information is playing a clearly in central role in what we would otherwise call our own internal states that make up part of our spirituality. It's the kind of search that we would do nowadays with a search engine.

Eckhart Tolle goes on to talk about "obedience to what came out of the present moment"...it guided his decision making process on where to be when. He goes on to say, "...don't do it on an impulse that is a restless impulse or comes out of any kind of negative emotion". If people listen to what he is saying, and a lot do, and if they combine their search for information with interaction with search engines, I land at the following conclusion: our individual spiritual development is not disconnected from our search engines and especially not from our experience of interacting with them.

In the end, the reason I blog about this might just be that I want to use the YouTube link to that Eckhart Tolle video that will take you right to the jump-in point that I am writing about
http://youtu.be/K1_R3uKJOB4?t=4m18s Goodness knows how much time I've spent discussing video fragment linking and trying to get research money to work on it as a searching speech problem -- I really get a kick out of being able to link into the stream.

We are late releasing the data for the MediaEval 2011 Genre Tagging task. The initial delay was small, but then other things just got in the way compounding the situation. Today, I am trying to be very present in the moment, in order to ignore the stress that I feel about being so late and be very careful about getting the release right the first time around.


And today's experience reminds me of how careful we need to be in all our research. If our search engines are part of our spiritual worlds, we need to design our algorithms and applications with awareness of their potential impact on the trails that we following in our paths of personal development and on the collective, common digestive system of humanity.

[1] Besser, J., Larson, M., Hofmann, K., Podcast Search: User Goals and Retrieval Technologies, Online Information Review: The international journal of digital information research and use, Vol. 34, No. 3, pp. 395-419, 2010.

Saturday, June 4, 2011

LikeLines: Crowdsourced intelligent mulitmedia player

Today we're doing the putting the final touches on LikeLines: Video highlights via web-scale aggregation of moments that viewers like, our entry to the mozilla Drumbeat Unlocking Video challenge, which is closing tomorrow.

The challenge addresses the question "How can new web video tools transform news storytelling?" Our answer to this question is the paradigm of distributed directing that allows news reports to be generated automatically, but without a central reporter. The raw material is footage captured by individuals with cameras and mobile phones who witness an event. One challenge faced is how to filter this footage: in particular, how to find the most interesting points? LikeLines gives the answer to this question.

The LikeLines concept is basically a heatmap that shows how many people found certain portions of a video interesting. You use it if you don't want to watch a video all the way through. Instead, you click the heatmap to jump in to just to the places that are worth watching start watching from there. What's worth watching is decided on the basis of what other viewers found worth watching -- either they tag those segments explicitly by clicking a "like" button or else they let the player record their stop, starting and cuing behavior. We also want the player to be able to make use of multimedia content analysis (visual analysis or speech recognition) in order to be able to "seed" interesting moments. This sort of seeding user contributions with multimedia content analysis has been used by our colleagues:

Ewine Smits and Alan Hanjalic. 2010. A System Concept for Socially Enriched Access to Soccer Video Collections. IEEE MultiMedia 17, 4 (October 2010), 26-35.

Our entry is in the form of video:



The video has been finished for a while now, and now I am just adding some text to make it clear that the idea is elegant, but also quite clever in that you combine user input and multimedia content analysis, which allows you to bootstrap from raw video.

It's a little crazy trying to write, because I need to switch out of research paper mode in to the mode of "hey, this will really work" and "hey look everyone, this is totally needed, totally non-trivial and totally does not exist anywhere yet". I am working now (when I stopped to write a blog post) on a sentence communicating that we can address the cold start issue with content analysis based seeding. And that verification using content analysis will help to control spam. And that if all goes well the whole thing should be able to learn by itself: It will require some R&D effort, but all the pieces of technology needed already exist.

Also, I had a little bit of trouble getting the right tone for the biography. So we're big shots at a cool technical university in the Netherlands? I guess that's important to communicate. But how to say that we are also passionate about supporting distributed and democratic news? Do I divulge that the first draft of LikeLines was churned out on a bus from Boston to Portland, the video was recorded in a long after-hours effort, and the whole thing has been discussed in every detail in chat sessions?

And how to communicate that we are doing this because it's what we love to do? We had some light-hearted lines in the bio to convey this tone (about our cat-video habit on YouTube and about me largely eschewing social media for the traditional postcard), but those got dumped in favor of some harder hitting facts about our experience in this area: right people, right skills, right place, right time...

In the end, I'm also in this to experience the crowdsourcing aspect of working on innovation in a open collaboration environment. What a breath of fresh air in the daily grind of publish and perish. And the giddy joy of communicating a concept that is ripe, feasible and useful.

Sunday, May 29, 2011

Tagging Love and Affection: Part II

Perhaps even the more interesting thing about the wedding in terms of modern media was the interaction between the professional photographers and the wedding guests who were taking pictures. It seemed like these were two completely different activities in terms of the results that they were aiming to produce. The photographers created an amazing album of storybook moments -- and the guests -- well, speaking for myself at least -- took pictures of people as people that they knew.

The professional photos actually included shots of members of the bridal party and guests photographing the bride and groom and each other. There was a particular dramatic one of the best man from the back taking a picture of the groom and you can see the groom twice: once over the shoulder of the best man in the display of his cell phone and once sitting as the main subject of the image.

There is another one in which one of the bridesmaids is taking a picture of the newly weds. It's like this act of greeting, an 'I was there with you in your moment of bliss and I was so so happy for you.' Taking a picture is like smiling, waving or winking at someone -- except that it is asynchronous, delayed in time. From the past, a shout out, "Congratulations!"

I was struck how the professional photographers were able to use the act of photo-taking as a way to depict the love and affection between friends and family members. Its not just the places that we tag with our photos, as I've discussed in a previous blog post, but its people, too.

And this is where the difference between the professional photographers and the guests really became apparent. I was using the little camera of my mom, so my pictures of the wedding are not qualitatively speaking very good. It would probably take some improvement of my photographic skills and not just a better camera to get high quality pictures.

But there was something that struck us about them. I naturally looked for the people in our family who we see in frequently and took pictures of people talking that only get to see each other once every several years -- if at all. I took pictures of people holding the youngest baby in our extended family -- that capture the moment that generations within the extended family meet for the first time. I took pictures that showed siblings engrossed in conversations with each other -- showing how the intensity of how we speak and how we listen. The photographers didn't know us and although the wedding pictures were beautiful, aesthetically not to be surpassed, we really love to look at the personal pictures, because they are somehow more "us".

Maybe the "real" pictures are the pictures that we take that mark love and affection. They encode our personalities, our common past and our hope for the future.

The implications are quite large for the field of multimedia retrieval, as revealed by the following line of reasoning: Life is finite. We only live so long and can support close relationships with so many people. If Dunbar is right it is a very limited number indeed. If we take photos at particular moments, such as moments of expression of affection that I am describing here, then the number of total pictures that we take is also limited. If we keep on insisting that our multimedia retrieval algorithms must be able to handle millions and millions of photos, then we run the danger of missing out on developing important techniques. Multimedia algorithms developed for relatively small numbers of photos can afford to be computationally more complex. If we ignore the "small set" problem, we run the danger of not developing the best possible algorithms to personal multimedia retrieval challenges.

As final comment, I can add that during the wedding I was already challenged by an image retrieval problem. I took maybe 200 photos. I wanted to show a special photo of my mom -- taken a few minutes back -- to the cousins I was sitting with at the dinner table. It took me so long to flip through the index to find those photos on that small screen. Very disruptive for dinner conversation.

I had the idea that I was the one that should have been wearing the GSR sensor. I am sure that my affective peak was physiologically measurable when I took the picture of my mom, saw it on my camera display for the first time and realized I had gotten a once-in-a-lifetime shot. If my camera display could take me right to peak pictures, it could be a much more functional device: transcending a capture to also support storytelling as well.

On second thought skip the GSR. I'm sure I jumped up and down. The accelerometer on a mobile phone could have picked that up. If all else fails, make a photo taking app that encourages me to shake the thing when I notice I like a picture. Better stop blogging and start implementing.

Tagging Love and Affection: Part I

Oh, my baby cousin is all grown up and got married. The wedding was a fairy tale: the kind that they show in the last scene of a good movie where you then sit through the entire credits in hopes that your eyes are reasonably dried up by the time you walk out into the bright public space beyond the theater.

The picture shows the bride's foot. Maybe one would expect a glass slipper, but this is a galvanic skin response sensor that recorded throughout the ceremony so that the bride and groom can later transcribe a mutual emotional trajectory. The sensor records skin conductance, which is related to moisture, and the signal must have been off the charts -- it was a wonderful wedding and but it was also a beautiful sunny day. Sunny in the intense sense, where you gain an practical understanding of why the British royal family feels that weddings require large-brimmed hats.

The nuptial couple must have been perspiring at least lightly and I suspect their respective sensors registered one continuous high throughout the proceedings. And indeed: I wish them this for their married lives that their joint signal continuously reflects a wonderful life experience. Or, if it's not a continuous high, may they at least be gifted with the ability to recognize any large dips as outliers and ignore them.

Lasting love is a complex phenomenon and I suspect that the most reliable expressions of it are more under our conscious control. In the next post I relate it to ... picture taking!