Thursday, October 20, 2011

Deep Link to Delft Technology Fellowship

Being educated in the US and being a scientist in Europe is sometimes quite tough. I need to continuously use a sort of filter that tells me that although I am hearing X, I need to pause and carefully consider and realize that the person is really saying Y. One particularly painful example, was unfortunately provided by our rector magnificus, the president of our university, in a recent interview. In promoting a new program to attract female scientists to the TU Delft, he said '...vrouwelijke wetenschappers zijn minstens zo talentvol als mannelijke wetenschappers.' which translates in English as 'female scientists are at least as talented as their male counterparts'. Ouch.

This statement does not work in the US academic context, because it fails gender symmetry. Gender symmetry can be diagnosed with the following test: flip the polarity of gender terms (e.g., 'woman', 'man', 'male', 'female') in a statement, and determine whether the resulting statement retains meaning within the context.

Let's try it. Flipping polarity of gender terms in his sentence yields, '...male scientists are at least as talented as their female counterparts'. This sentence is clearly interpretable, but no longer has a meaning that fits the context.

Contrast that with an alternate sentence such as: 'There is no discrepancy in talent between male and female scientists'. This sentence has the same declarative content, but it passes the gender symmetry test because you can substitute it with 'These is no discrepancy in talent between female and male scientists'.

Of course, in this case, a further problem arises. This sentence has the implicature that there is some reason for which this fact needs to be asserted in the first place. The act of pronouncing this sentence communicates that the speaker does not consider the point to be completely obvious, but rather feels that it needs to be explicitly asserted. One might choose against even this alternative sentence in order to avoid sending the message that one feels that there is someone out there that still needs to be convinced on the point of talent equivalence between male and female scientists. But on the whole, this alternative could be considered the 'best practices' formulation, should one indeed find oneself in a situation where it was necessary to make a statement comparing the relative scientific talent of men and women.

What my filter tells me is that although X was said in this case, what was meant is Y. And concerning Y, I rather suspect that our rector magnificus harbors the personal opinion that women have perhaps even a teensy bit more science talent than men and that in fact he is saying, "at least as (if not more) qualified". Whether or not that is true, it's safe to say that he is of the opinion that our university would, at this point in time, benefit from hiring additional women.

One of the research topics that I am interested in as a multimedia retrieval scientists is developing algorithms for the retrieval of jump in points (JIP) in video. JIPs allow the viewer to click directly to a certain relevant point in a video. On YouTube, they are called deep links. JIPs make it possible to share or to comment about particular points of a video, just as I am currently doing with this post. The deep link to the relevant section of the interview under discussion is the following:

http://youtu.be/wvto6MWXE6k?t=35s

The current status of technology on the Web is that it is possible to comment on JIPs or share them, but search engines don't return them as results. Together with colleagues within the Netherlands and across Europe I am developing and helping to promote the development of JIP retrieval in the MediaEval Rich Speech Retrieval task (see the feature on MediaEval 2011 in MMRecords for a brief description.) Such technology would allow search engines to return pointers to specific time points within video that are relevant to user queries.

At the end of the day, I am more interested in the scientific questions raised by the task of JIP multimedia retrieval than I am in the gender issue. Since grade school, I have frequently been the "only girl" involved in whatever activity fascinated me. You don't know it any other way, so you don't really notice. I contribute what I can to the discourse on promoting gender balance, not so much because of myself, but because I find it wasteful if I feel that women who I am mentoring are somehow holding themselves back.

When I first came to Delft, I contributed the following comment on improving the working climate at the Faculty of Electrical Engineering, Mathematics and Computer Science (EEMCS). This is the point of view that I still stand by so I include it here to complete my comment on the deep link.

Response on the 2009 Challenging Gender survey
The way of improving the working climate at EEMCS would be to address the gender imbalance within a larger program of promoting diversity into the Faculty of EEMCS. A faculty that includes international scientists addressing multi- and trans-disciplinary questions is automatically going to be more comfortable for women, since gender differences become just one of many differences of background and perspective that make the faculty richer and more productive.

Any effort invested in promoting inclusion of scientists/researchers that have pursued non-traditional career tracks (e.g., completing their PhD at an older or younger age, taking time off, switching disciplines mid-career) will automatically make women feel more welcome. When women feel welcome, they will also feel confident that the effort that they invest will be rewarded by a long and productive career in the EEMCS, establishing a virtuous cycle.

Everyone benefits from the promotion of diversity. For example, in this kind of climate, a researcher who has worked in the faculty for years will feel more comfortable about taking the risk of investigating a new class of algorithms or applying expertise accumulated in one domain to solving a problem in a radically different domain.

Positive side-effect: If everyone benefits, then women will not be burdened by the (perceived) need to fight the prejudice that they have been hired due to their gender and not due to their competence.

By promoting diversity, both in terms of scientific expertise and also in terms of other characteristics (cultural, religious, linguistic, socio-economic, sexual orientation as well as gender), the faculty will draw on a larger pool of talent and increase its productivity and capacity for creation and invention.

Working at TU-Delft, you see "Challenge the future" written everywhere...sometimes in unexpected places. As a woman this speaks to me in a special way: it says that the future at the TU-Delft is not set up to be carbon copy of the past. Because of the "challenge the future" attitude, I have confidence that the demographics of my department will shift naturally as we the Faculty of EEMCS continues to mature, extend and innovate scientifically.

Wednesday, September 28, 2011

Search Computing and Social Media Workshop

Today, in Torino, Italy, was the day of the Search Computing and Social Media Workshop organized by Chorus+, Glocal and PetaMedia. Being the PetaMedia organizer, I had the honor of opening the workshop with a few words. I tried to set the tone by making the point that information is inherently social, being created by people, for people. Digital media simply extends the reach of information, letting us exchange with others and with ourselves over the constraints of time and space.

The panel at the end of the day looped back around to this idea to discuss the human factor in search computing. We collected points from the workshop participants on pieces of paper to provide the basis for group discussion. I made some notes about how this discussion unrolled. I'm recording them here while they are still fresh in my head.

We started by tackling a big, unsolved issue: Privacy. The point was made that the very reason why social media even exists is that people seem driven in some way to give up their privacy, share things about themselves that no one would know unless they were revealed. Whether or not users do or should compromise their own privacy by sharing personal media was noted to depend on the situation. For some people it's simply, obviously the right thing to do. Concerns were raised about people not knowing the consequences: maybe effectively I am a totally different person five years from now than I am now. But I am still followed by the consequences of today's sharing habits. In the end, the point was made that if the willingness to among users to share stops, we as social media researchers have not much else left to examine.

Next we moved to the question of events in social media: Human's don't agree about what constitutes and event. Wouldn't it just be easier to just adopt as our idea of an event whatever our automatic methods tell us is an event? Effectively we do this anyway. We have no universal definition of an event. There may be some common understanding or conventions within a community that define what an event is. However, these do not necessarily involve widespread consensus: they may be personal and they may evolve with time. For example, the event of "freedom"? Most people agreed that freedom was not an event.

An event is a context. That's it. At the root of things, there are no events. Instead, we use concepts to build from meaning to situational meaning -- to the interpretation of the meaning of the context. Via this interpretation, the impression of event emerges. In the end, meaning is negotiated.

If we say events are nothing, we wouldn't be able to recognize them. Or, does the computer simply play a role in the negotiation game. The systems we build "teach" us their language and we adapt ourselves to their limitations and to the interpretative opportunities that they offer.

Then the question came up about the problems that we choose to tackle as researcher. "Are we hunting turtles because we can't catch hares?" This bothered me a bit, because assuming you can easily catch a turtle, they are quite difficult to kill because of the shell. The hare would be easier. Do our data sets really allow us to tackle "the problem"? The question presupposes that we know what "the problem" is, which may be the same as solving the problem in the first place. Maybe if we can offer the user in a give context enough results that are good enough, they will be able to pick the one that solves "the problem". Perhaps that's all there is to it. Under such an interpretation, the human factor becomes an integral part of the search problem.

In the end, a clear voice with a succinct take home message:
How can we efficiently combine both the human factor and technology approaches?
"The machine can propose and the user can decide."

The discussion ended naturally with a Tim Berners Lee quote, reminding us of the original intent of social effect underlying the Web. We adjourned for some more social networking among ourselves, reassuring ourselves that as long as we were still asking the question we shouldn't expect to find ourselves completely off track.

Friday, September 2, 2011

MediaEval 2011: Reflections on community-powered benchmarking

The 2011 season of the MediaEval benchmark culminated with the MediaEval 2011 workshop that was held 1-2 September in Pisa, Italy at Santa Croce in Fossabanda. The workshop was an official satellite event of Interspeech 2011.

For me, it was an amazing experience. So many people worked so hard to organize the tasks, to develop algorithms and also to write their working notes papers and prepare their workshop presentations. I ran around like crazy worrying about logistics details, but every time I stopped for a moment I was immediately caught up in amazement of learning something new. Or of realizing that someone had pushed a step further on an issue where I had been blocked in my own thinking. There's a real sense of traction -- the wheels are connected with the road and we are moving forward.

I make lists of points that are designed to fit on a Power Point slide and to succinctly convey what MediaEval actually is. My most recently version of this slide states that MediaEval is:
  • ...a multimedia benchmarking initiative.
  • ...evaluates new algorithms for multimedia access and retrieval.
  • ...emphasizes the "multi" in multimedia: speech, audio, visual content, tags, users, context.
  • ...innovates new tasks and techniques focusing on the human and social aspects of multimedia content.
  • ...is open for participation from the research community
I make these lists and they capture the external reality of what we do, but actually I have no real understanding of how MediaEval works -- of how exactly the traction arises.

At the workshop I attempted to explain it with a bunch of circles drawn on a flip chart (image above). The circles represent people and/or teams in the community. A year of MediaEval consists of a set of relatively autonomous tasks, each with their own organizers. Starting in 2011, we also required that each task have five core participants who commit to crossing the finishing line on the tasks. Effectively, the core participants started playing the role of "sub-organizers", supporting the organizers by doing things like beta testing evaluation scripts.

This set up served to distribute the work and the responsibility over an even wider base of the MediaEval community. Although I do not know exactly how MediaEval works, I have the impression that this distribution is a key factor. I am interested to see how this configuration develops further next year.

MediaEval has the ambitious aim of quantitatively evaluating algorithms that have been developed at different research sites. We would like to determine the most effective methods for approaching multimedia access and retrieval tasks. At the same time, we would like to retain other information about our experience. It is critical that we do not reduce a year of a MediaEval task to a pair (winner, score). Rather, we would like to know which new approaches show promise. We would like to know this independently of whether they are already far enough along in order to show improvement in a quantitative evaluation score. In this way, we hope that our benchmark will encourage and not repress innovation.

I turned from trying to understand MediaEval as a whole to trying to understand what I do. Among all the circles on this flip chart, I am one of the circles. I am a task organizer, a participant (time permitting) and also play a global glue function: coordinating the logistics.

The MediaEval 2012 season kicks-off with one of the largest logistics tasks: collecting people's proposals for new MediaEval tasks, making sure that they include all the necessary information, a good set of sub-question and getting them packed into the MediaEval survey. It is on the basis of this survey that we decide the tasks that will run in the next year. We use the experience, knowledge and preferences of the community in order to select the most interesting, most viable tasks to run in the next year and also to decide on some of the details of their design.

Five years ago, if someone told me I would be editing surveys for the sake of advancing science, I would have said they were crazy. Oh, I guess I also ordered the "mediaeval multimedia benchmark" T-Shirts. That's just what my little circle in the network does.

Let's keep moving forward and find out where our traction lets us go.

Thursday, August 25, 2011

LinkedIn does a SlippedIn: And LinkedIn users themselves do damage control
















Like me, you possibly had to have someone else bring to your attention the new social advertising functionality of LinkedIn: And the fact that "on" was introduced as the default setting. Here's how to turn it off:
http://www.julianevansblog.com/2011/08/how-to-manage-your-linkedin-social-advertising-privacy.html

I got an email this morning from someone close to me, S., whose colleague, C., had sent them a message with a Dutch translation of these directions on how to turn the social advertising off. S. declared happily, "The community is really strong". LinkedIn pulled a now-classic social network move and the community moves to push back against it. If there wasn't a name for it already, we can now conveniently refer to it as a SlippedIn.

SlippedIn or slipped up? The fact that this changed behind my back really makes me angry at LinkedIn: Are they going to lose their community?

Well, no. Because actually in sending this mail C. is engaging, probably without her conscious knowledge, in the ultimate form of social advertising. By alerting us to the problem and letting us know how to fix it, C. is mediating between LinkedIn and the community that uses the LinkedIn platform. She is making it possible for all of us to be really p.o.ed at LinkedIn, but still not leave the LinkedIn network because we have the feeling that our community itself has created the solution that keeps us in control of our personal information.

S.'s attitude "The community is really strong" is natural. Because C. caught this feature being slipped in and let us know how to fight it, we now have the impression that we somehow have the power to band together and resist the erosion of the functionality that we signed up for when we joined LinkedIn. C.'s actions give us the impression that although what LinkedIn did is not ok, that LinkedIn is still an tolerable place to social network because we have friends there and that we are in control and can work it out together.

C. has really be used. She is unwitting broadcasting in her social circle a sense of security that everything will be all right. We completely overlook the point that we have no idea of what goes on beyond the scenes that might go on unnoticed by C. or the other C.-like people in the network. We are given the false impression, that whatever LinkedIn does that we find intolerable, that we will be able to notice it and work together to fix it.

We cannot forget that LinkedIn is a monolithic entity: they write the software, they control the servers. What ever feeling that we have that we can influence what is going on is supported only by our own human nature to simply trust that our friends will take care of us. LinkedIn is exploiting that trust to create a force of advocacy for their platform as they pursue a policy aimed at eroding our individual privacy.

Last week I spent a great deal of time last week writing on a proposal called "XNets". Basically, we're looking for a million Euros to help develop robust and productive networking technology that will help ensure that social networking unfolds to meet its full potential. Our vision is distributed social networking: let users build a social network platform where there is no central entity calling the shots.

However, it's not just the distributed system that we need it is the consciousness. I turned the social advertising functionality off and have for the moment the feeling that it is "fixed". But getting this fixed was not C.'s job. C. is not all-seeing nor can she help her friends protect themselves against all possible future SlippedIns. C. should not be doing damage control for LinkedIn. We the community are strong, but we are not omnipotent. The ultimate responsibility for safe-guarding our personal data lies with LinkedIn itself.

Saturday, August 13, 2011

Human computational semantics

What I termed "Human computational relevance" in my previous blog post is probably more appropriately termed "Human computational semantics". The model in the figure in that post can be extended in a straightforward manner to accommodate "Human computational semantics". The model involves comparing multimedia items (again within a specific functional context and a specific demographic) and assigning them a pair-wise similarity value according to the proportion of human subjects that agree that they are similar.


Fig. 1: The similarity between two multimedia items is measured in terms of the the proportion of human subjects within a real-world functional context and drawn from a well-defined demographic that agree that they are similar. I claim that this is the only notion of semantic similarity that we need.

I hit the ceiling when I hear people describe multimedia items as "obviously related" or "clearly semantically similar". The notion of "obvious" is necessarily defined with respect to a perceiver. If you want to say "obvious", you must necessarily specify the assumption you make about "obvious to whom". Likewise, there is no ultimate notion of "similarity" that is floating around out there for everyone to access. If you want to say "similar", you must specify the assumption that you make about "similar in what context."

If you don't make these specifications, then you are sweeping an implicit assumption you are making right under the rug and it's sure to give you trouble later. It's dangerous to let ourselves lose sight of our unconscious assumptions of who our users are and what the functional context actually is in which we expect our algorithms to operate. Even if it is difficult to come up with a formal definition at least we can remind ourselves how slippery these notions are be. It seems that we naturally as humans like to emphasize universality and our own commonality, and that in most situations it's difficult to really convince people that "obvious to everyone" and "always similar" are not sufficiently formalized characterizations to be useful in multimedia research. However, in the case of multimedia content analysis the risks are too great and I feel obliged to at least try.

A common objection to the proposed model runs as follows: "So then you have a semantic system that consists of pairwise comparisons between elements, what about the global system?" My answer is: The model gives you local, example-based semantics. The global properties emerge from local interactions in the system. We do no require the system to be globally consistent, instead we gather pairwise comparisons until a useful level of consistency emerges.

Our insistence on a global semantics, I maintain, is a throwback to the days that we only had conventional books to store knowledge. Paper books are necessarily linear, necessarily of a restricted length and have no random access function. So, we began abstracting and organizing and ordering to back human understanding of the world into an encyclopedic or dictionary form. It's a fun and rewarding activity to construct compendiums of what we know. However, there is no a priori reason why a semantic system based on a global semantic model must necessarily be chosen for use by a search engine.

Language itself is quite naturally defined as a set of conventions that arise and are maintained via highly local acts of communication within a human population. Under this view, we can ask about Fig. 1, why I didn't draw in connections between the human subjects in order to indicate that the basis of their judgements rests in a common understanding -- a language pact as it were. This understanding is negotiated over years of interaction in a world that it exists beyond the immediate moment at which they are asked to answer the question. Our impression that we need an a prior global semantics arises from the fact that there is no practical way to integrate models language evolution or personal language variation into our system. Again, it's sort of comforting to see that when people think about these issues their first response is to emphasize universality and our human commonality.

It's going to hurt us a little inside to work with systems that represent meaning in a distributed, pairwise fashion. It goes against our feeling, perhaps, that everyone should listen to and understand everything we say. We might not want to think too hard about how our web search engines have actually already been using a form of ad hoc distributed semantics for years.

In closing: The model is there. The wider implications of its existence are that we should direct our efforts to solving the engineering and design problems necessary to be able to efficiently and economically generate estimations of human computational relevance and also of the reliability of these estimates. If we accomplish this task, we are in a position to be able to create better algorithms for our systems. Because we are using crowdsourcing -- computation carried out by individual humans -- we also need to address the ethics question: Can we generate such models without tipping the equilibrium of the crowdsroucing-universe so that it disadvantages (or fail to advantages) already fragile human populations?

This post is dedicated to my colleague David Tax: One of the perks of my job is an office on the floor with the guys from the Pattern Recognition Lab -- and one of the downsides is a low-level, but nagging sense of regret that we don't meet at the coffee machine and talk more often. This post articulates the larger story that I'd like to tell you.

Subjectivity vs. Objectivity in Multimedia Indexing

In the field of multimedia, we spend so much time in discussions about semantic annotations (such as tags, or concept labels used for automatic concept detection) and whether they are objective or subjective. Usually the discourse runs along the lines of "Objective metadata is worth our effort, subjective metadata is too personal to either predict or be useful." Somehow the underlying assumption in these discussions is that we all have access to an a priori understanding of the distinction between "subjective" and "objective" and that this distinction is of some specific relevance to our field of research.

My position is that, as engineers building multimedia search engines, if we want to distinguish between subjective and objective we should do so using a model. We should avoid listening to our individual gut feelings on the issue (or wasting time talking about them). Instead, we should adopt a the more modern notion of "human computational relevance" which, since the rise of crowdsourcing, has entered into conceivable reach.

The underlying model is simple: Given a definition of a demographic that can be used to select a set of human subjects and a definition of a functional context in the real world inhabited by those subjects, the level of subjectivity or objectivity of an individual label is defined as the percentage of of human subjects who would say "yes, that label belongs with that multimedia item". The model can be visualized as follows:

Fig. 1: The relevance of a tag to an object is defined as the proportion of human subjects (pictured as circles) within a real-world functional context and drawn from a well-defined demographic that agree on a tag. I claim that this is the only notion of the objective/subjective distinction relevant for our work in developing multimedia search engines.

Under this view of the world, the distinction between subjective and objective reduces to the inter-annotator agreement under controlled conditions. I maintain that the level of inter-annotator agreement will also reflect the usefulness that the tag will have deployed within a multimedia search engine designed for use within the domain defined by the functional context by the people in the demographic. If we want to assimilate personalized multimedia search into this picture we can define it within a functional context for a demographic consisting only of one person.

This model reduces the subjective/objective difference to a estimation of the utility of a particular annotation within the system. The discussions we should be spending our time on are the ones about how to tackle the daunting task of implementing this model so as to generate a reliable estimates of human computational relevance.

As mentioned above, the model is intended to be implemented on a crowdsourcing platform that will produce an estimate of the relevance of each label for each multimedia item. I am as deeply involved as I am with crowdsourcing HIT design because am trying to find a principled manner to constrain worker pools with regard to demographic specifications and with regard to the specifications of a real-world function for multimedia objects. At the same time, we need useful estimators of the extent to which the worker pool deviates from the idealized conditions.

These are daunting tasks and will, without doubt, require well-motivated simplifications of the model. It should be clear that I don't claim that the model makes things suddenly 'easy'. However, it is clearly a more principled manner of moving forward than debate on the subjectivity vs. objectivity difference.

Continued...

Monday, August 8, 2011

The power of collaborative competition

LikeLines is a crowdsourced intelligent mulitmedia player that aggregates viewer clicks along a video timeline to create a heatmap of where viewers find a video most interesting.



Today was the day that Raynor submitted his final project for the Knight-Mozilla learning lab. The Learning Lab is part of the Knight-Mozilla News Technology Partnership (shortened to MoJo) that is run as a contest. I was a bystander and watched Raynor go through the process of attending lectures online, writing blogposts and exchanging comments and tweets with other members of the lab.

I was just amazed at the people involved in this contest: in their ability to develop their own idea and distinguish themselves, but at the same time support each other and collaborate as a community. It's nice to talk about crowdsourced innovation, but it's breathtaking to experience it in action.

The results are reflected in how far LikeLines has come since when I first posted on it at the beginning of June. Raynor looked at me one day and said, "It's an API"...and we realized that this is not just an intelligent video player it is a whole new paradigm for collecting user feedback that can be applied in an entire range of use cases.

From one day to the next we started talking about time-code specific video popularity, which we quickly shorted to "heatmap metadata".

Whatever happens next, whether Raynor proceeds to the next round, I already have an overpowering sense of having "won" at MoJo. It really solidified my belief in the power of collaborative competition as a source of innovation -- and a force for good.

I am an organizer in the MediaEval benchmark and this is the sort of effect that we aspire to: bringing people together to pull towards a common goal simultaneously as individuals and as a community.

There needs to be a multiplicity of such efforts: they should support and learn from each other. I can only encourage the students in our lab to get out there and get involved, both as participants and as organizers.

One day last week we were in the elevator heading down to lunch and Yue Shi turned to me and said. Do you realize that of the people standing in the elevator, there are five PhD students submitting entries in five different competitions?
True to usual style, my first reaction is, "Hey people, what happened to TRECVID?" We are also make an honest effort to submit to TRECVID this year. I watched that happen...and then not happen.

But then I gave myself permission, there in the elevator to turn off the bookkeeping/managing mechanism mechanism in my head -- and just go with my underlying feeling of what we were doing as a lab. It's the feeling of wow. Everybody doing their own thing, but at the same time being part of this amazing collaborative competitive community.

The elevator doors opened and as we passed through I thought, it seems like the normal daily ride that we're taking, but when you look a bit deeper you can see the world changing and how the people in my lab pool efforts to change it.