An important innovation of MediaEval 2012 was the "Brave New Task" track. MediaEval is a multimedia benchmark that offers and promotes challenging new tasks in the area of multimedia access and retrieval. We focus on tasks that emphasize multiple modalities ("The 'multi' in multimedia") and that have social and human relevance.
Brave New Tasks were introduced because we noticed that there is rather a tension between benchmark evaluation and innovation. Benchmarking is essentially a conservative activity: we want to compare algorithms on the same task using the same data set and the same evaluation procedure. This sameness allows us to chart progress with respect to the state of the art, especially over the course of time. How do we innovate, when the key strength of benchmarking is that we repeatedly do the same thing?
We innovate by tackling new problems. However, in order to create a successful benchmarking task from a new problem, a number of questions must be answered. Is the problem suitable for evaluation in a benchmark setting? What sort of data is needed to evaluate solutions developed by benchmark participants? How much effort is needed to create ground truth? Do we need to refine our definition of the task and of the evaluation procedure? Is there an actual chance that algorithms can be developed to solve the task and what resources are needed? Is there a critical mass of interest in solving this problem? Are the solutions appropriate for application?
The easy way forward would be to insist that there are clear answers to all of these questions prior to running a task. In some cases, is will be possible to gather the answers. In others, however, it will not. Forcing tasks to have answers before attempting to create a benchmark poses a serious risk that researchers will avoid the truly challenging and innovative tasks because they receive the message that they need to "play it safe."
People in the MediaEval community rattle their swords and shields when they are told that they need to "play it safe." Brave New Tasks support innovation in MediaEval by
incubating tasks in their first year, allowing the task organizers to
answer these questions. We value the advantages that the conservative aspects of benchmarking bring to the community, but we also thrive by taking risks. The Brave New Task track is a lightly protected space that allows us to take the risks that allow our benchmark to continue to renew itself.
Because people have asked me about Brave New Tasks in MediaEval 2013 and "How did you do it?" I am providing here a more detailed description of how it works and how we anticipate that it will develop in 2013:
To start, let me write a few words about makes a main "mainstream" task in MediaEval (i.e., a task that is not a Brave New Task). At the end of the calendar year, MediaEval solicits proposals from teams who are interested in organizing a task in the next MediaEval season. For an example, see the MediaEval 2013 call for task proposals. Whether a proposal is accepted as a MediaEval task depends on the interest expressed on the MediaEval survey. The survey is published in the first days of January and circulated widely to the larger research community.
During the survey, task proposers gather information on who is interested in carrying out their tasks. By the time the survey concludes, the proposers must have promises from five core participants (who are not themselves organizers) who will cross the finish line of the task (including submitting, results writing the working notes paper and attending the MediaEval workshop) come "hell or high water". This selection criteria is set up so that we have a minimum number of results to compare across sites for any given task---if there are only one or two, we don't get the "benchmark effect".
Tasks that the community finds interesting and promising, but that do not necessarily meet these stringent selection criteria, can be selected as Brave New Tasks. The difference between a Brave New Task and other MediaEval tasks is that these tasks are new, and ideally also a scientifically risky (in the responsible sense of "risky").
Brave New Tasks are run "by invitation only". The "invitation only" clause does not make the task exclusive: anyone who asks the task organizers can be granted an invitation. Instead, the clause allows the tasks to handle unexpected situation by, if necessary, decoupling their schedules from the main task schedules to accommodate unforeseen delays in data set releases. Participants of past editions of MediaEval will recognize the usefulness of a mechanism that makes the benchmark robust to unexpected situations.
Further, Brave New Tasks do not require their participants to submit working notes papers or attend the workshop. The "only" requirement that the task must fulfill is to contribute an overview paper in the MediaEval working notes proceedings that sums up the task and presents and outlook for future years. One or more of the organizers attends the workshop to make the presentation and participate in the discussion about whether the task should target developing into a mainstream task in the next year.
A "Brave New Task" is encouraged to go far beyond the minimum requirement. In fact, 2012 saw one of the Brave New Tasks "Search and Hyperlinking" achieve the scope of a mainstream task, with six working notes papers from task participants appearing in the MediaEval 2012 working notes proceedings. The task was effectively indistinguishable from mainstream tasks in its contribution to the benchmark.
In 2013, we plan to strengthen the Brave New Task track by providing them with more central support. The tasks will be run under the same infrastructure as the mainstream tasks and decoupled from the schedule only if it is absolutely necessary. They will also be given the option of using the central registration system.
Brave New Tasks have been a successful innovation in MediaEval 2012 and one that we hope to strengthen in the future. I'd like to end by pointing out that it is not so much the "rules" of Brave New Tasks that have made them such a success, but rather the efforts of the Brave New Task organizers. Success is dependent on having a group of devoted researchers with a vision for a new task idea and the capacity and stamina to see it through the first year...including long hours spent reading related work, developing new evaluation metrics (if necessary), contacting and following up with participants, collecting data and creating ground truth. It is not so much the tasks themselves that are brave, but the organizers who are fearless and relentless in their pursuit of innovation.
Forward, charge!
Wednesday, January 2, 2013
Friday, November 30, 2012
ImageNet and the Edge of the World: On visual concept labels for images
![]() |
| Google Image Search results for the query "two-year-old horse" |
ImageNet is very cool. If I were a kid, this would have been better than any of those other picture dictionaries...it is fun just to click through and explore what exists in the world.
I recently made a video about this paper:
Jia Deng; Wei Dong; Socher, R.; Li-Jia Li; Kai Li; Li Fei-Fei; , "ImageNet: A large-scale hierarchical image database," Proc. IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2009. pp.248-255.
The video's down below, in case you want to get a quick overview of the paper...but in the video I am mostly focusing on discussing the crowdsourcing methods applied to create ImageNet.
Here, I will discuss "Edge of the ImageNet Image World". I identify the edge with the idea of a concept being "difficult to be illustrated". This concept I found mentioned in footnote 1 of the paper:
About 20% of the synsets have very few images, because either there are very few web images available, e.g. “vespertilian bat”, or the synset by definition is difficult to be illustrated by images, e.g. "two-year-old horse". (p. 249)
I would like to make the point that we may be radically underestimating the importance of "difficult to be illustrated" visual concepts in our multimedia information indexing systems.
Difficult to be illustrated? It is quite obvious that it is not inherently difficult to have a picture of a two-year-old horse.
I have a horse, two years ago I watched it being born, and I take a picture of it. Finished.
What is difficult is to find a group of people (annotators or users) who will look only at the picture (knowing nothing about me and the horse) and agree that the horse is two years old.
It is difficult for two reasons:
- Context of use: The concept "two-year old horse" is difficult to pin down exactly. Does a horse that is two-years and one day old still count as a two-year-old horse? It depends on what you are using the picture for. If you are using it for a collection of "horses on their second birthdays" it won't count. However, if you are using it to illustrate horses that are less than full grown, that day doesn't matter.
- Background of user: You have to know something about horses to distinguish a horse that is a foal (under one year) from one that is a colt or a filly (which Wikipedia tells us are terms that may be used until the horse is 3 or 4).
As multimedia researchers, we seem to assume that these "difficult to illustrate concepts" represent some marginal part of multimedia meaning. I mean, it's less that 20% of the concepts in WordNet that have this problem, so isn't it a good first approximation to just ignore them and focus on the 80% that are easily illustrated by images?
Context of use: We can just concentrate on the formal definitions of concepts. It's about delivering precise results lists when we search for images isn't it? Under that view, we can solve point 1. by deciding to use the most restrictive definition possible: the horse that turned three yesterday is no longer a two-year-old horse.
OK. So we all are totally annoyed at the guy who just sits immobile when we say, "Hand me that red screwdriver?" You climb down from the ladder just to hear him say, "I see a crimson screwdriver, but no red screwdriver." We are annoyed because we know that language is built to be used, and part of that use is the fact that we accommodate the meanings of words within their contexts of use.
But we learn to live with it. We realize the guy is literally right, so we grab the screwdriver ourselves and climb back up the ladder. We could learn to live with image search engines that behave like that as well, couldn't we?
Background of user: We can just concentrate on what the "man on the street" thinks about the image. It's about delivering results that are generally recognizable and not results that require some expert insight, isn't it? Under that view we can solve point 2. by deciding to use what a member of the general public would say about the image: it's a horse, probably not a grown up horse, but there's no telling if it's two years old.
Whoa. Hold your horses right there! Who gets to then decide who constitutes the "man on the street" of the "general public"?
Many people that I meet on the street in my daily life are not going to know the difference between quite obvious concepts like "bananas" vs. "plantains". It depends on what street I chose to look at.
With respect to many streets in Western Europe, "plantain" would be "difficult to be illustrated": people that can identify them are somehow considered experts. Not so in West Africa.
Irresponsible intuitions: In a split second, we as multimedia researchers can make a decision that seems "obvious", but that on closer consideration has potential to come back and haunt us.
We are reinforced to make these "obvious" decisions because they are the ones that allow us to continue on with our research with a minimal investment of resources in creating labeled image sets.
If I use restrictive, formal decisions, I don't have to turn to actual users of image search engines to try to understand how the "language of concepts" that they use when they search.
I also don't have to try to dig down to more subtle forms of cultural bias that exist in WordNet. Who of us has time to read a volume on cultural bias in dictionaries with contributions from 40 scholars?
In the end, although "difficult to be illustrated concepts" may constitute 20% of the concepts in WordNet, we have no idea of what percent of actual user image need might be related to these concepts. It could be huge!
Edge of the world: Google somehow gets it right. The search results at the top are returned by Google Images in response to the query "two-year-old horse". The first image occurs on the Internet in conjunction with the text "2 Year old Buckskin Quarter horse Colt". Someone apparently took a picture of their two-year-old horse and that seems to be right.
In the next picture, it's the kid and not the horse that's two, but that's pretty obviously wrong, and even amusing.
At the very least, this discussion allows us that to conclude that if ImageNet covers "The Image World", that is a very flat world indeed. It is easy to follow a "difficult to be illustrated" concept to the end of that world and stand there looking over the edge...
...ImageNet is a valuable research tool and serves the community well. However, we should all be aware of exactly where the edge of the ImageNet world is, not that we want to avoid it, but perhaps because that is exactly the place from which we want to leap off.
Labels:
cultural bias,
image annotation,
image search,
ImageNet,
meaning,
visual concepts,
WordNet
Saturday, August 25, 2012
Gender in Advertising Images: The Devil is in the Detail
The most obvious "bug" is the choice to include in the advertisement an image of a person. Since this misstep is a useful illustration of the limitations of visual depictions in multimedia, I decided to dedicate a blogpost to discussing it.
At first consideration, it seems obvious that our university should advertise using image of people. One of the reasons that I love working at TU Delft is the emphasis on solving societal problems. Using pictures containing people and not just technology wherever possible seems to be a good strategy for getting the importance of our work to address human and social challenges across.
However, a major limitation for visual depictions such as images and videos is "the curse of instance depiction". Basically, it is impossible to create such visual imagery without committing yourself to depicting a full range of details. You can't get across and abstract concept, for example, "car" without actually committing yourself to an instance of a single car existing in the real world, which you take to stand for all cars. Instead, you are going to need to show in your image a specific type, make and model.
Here, the concept that the ad is trying to convey is "professor". The "type, make and model" chosen to convey this concept are an adult of a certain gender and a certain age group, wearing glasses. It seems plausible that the person designing the ad was aware of the problem of instance depiction. The decision to use a model with a shaved head makes it possible to avoid depicting the hair color, which could serve to further specify the ethnic background or the age.
However, it is extremely difficult, if not impossible, to "hedge" on the gender question in images of the real world. A person depicted in a daytime work setting will generally be identifiable as a male person or a female person.
If we assume that the process by which we choose and interpret images that are being used to represent categories follows prototype theory, then the choice of a male to represent a TU Delft professor is no just unbalancing for the reader of the advertisement, but is very serious indeed. Prototype theory tells us that in our cognitive representations, some members of conceptual categories are more salient than others. We think of them first when we think of a category and we react to them more quickly when confirming category membership.
The use of a male person in this advertisement sends the message that males are the canonical professors at the TU Delft. Although men are clearly in the majority in the faculty, there is not any sort of a conscious intention at the university to keep the situation that way. In fact, I have the impression that everyone is working to shift their idea of how can be a professor to encompass a diverse demographic more directly representative of the general population.
Visual depictions in multimedia, i.e., images depicting the real world, are limited in what they can express because they deprive us of the possibilities of leaving certain details unpecified. What we have is a reversal of the saying "A picture is worth a thousand words." Instead, the spoken or the written word is able to express more in this case because human language can directly convey concepts without having to make use of specific instances to do so. In effect, the possibility for ambiguity or underspecification is makes human language more expressive that multimedia.
And so, the saying "The devil is in the detail" takes on a new shade of meaning.
What to do about the advertisement? I advise having a closer look at some advertising guidelines. Advertising Standards Canada, a non-profit self-regulation body for advertising, has a helpful list of guidelines for balancing gender representations in advertising online and surely Europe has a similar set of guidelines.
An "quick and dirty" solution is to look to see how other universities advertise. In the Economist, a general tendency to avoid imagery is readily apparent. For example, next to the TU Delft advertisement is a classical advertisement for Harvard faculty positions, whose only graphic content in the Harvard Business School logo.
I was cheered up again when my Google Googles app confirmed for me that the logo used was from the business school (i.e., distinct from the main Harvard Logo). It is my first use of Google Goggles for something other than just playing around with while hanging out with my multimedia information retrieval colleagues.
For completeness, I note a less obvious bug. The advertisement contains the text "Maximum employment: 38 hours per week (1 FTE)" In order to interpret this text, you need to know that "FTE" stands for "full time equivalent". 1 FTE means this position is a full time job. Contrary to what the text implies, no one the Safety Science processor to working 38 hours a week.
Labels:
advertising,
Economist,
gender,
Google googles,
images,
TU Delft
Wednesday, August 8, 2012
Worry-Free Social Sharing for Social Networks
| Flickr: Phil Wiffen |
In contrast to smoking, social sharing done right actually helps rather than hurts. In fact, the rise of online social networking and social multimedia sharing has been downright amazing technological development. Moved by this awe, last year in a project proposal, I effused that social networks are, "...a virtual prosthetic that extends the strong fabric of social connectivity critical to the well-being and growth of human societies into the online realm."
That proposal developed the idea of "worry-free social sharing": a social sharing client that would gently alert us when our sharing actions, in ways we do not intend, threaten to compromise our privacy---and then suggest alternative actions, which allow us to share our personal experiences, but in a wiser way.
Yes, people are responsible for their own actions. But in some cases, we as individual users do not have the understanding of multimedia analysis technology, or of the power of algorithms to combine different sorts of data to reveal facts about us that we thought were hidden. We all would need such understanding in order to allow us to make informed decisions about which types of social sharing is harmless and which types should better be avoided.
Even for the most savvy of us there are always surprises: Did you know that if you upload a video to YouTube and you carefully avoid geo-tagging it, but if you happen to be in a city and capture an ambulance siren in the background, that siren will serve to indicate in which city you are? Check out the work on multimodal location estimation [1]. Maybe you don't care if the world knows where you are, but if you do happen to be worried about having left your house empty during your vacation, it would be good to know that you just about betrayed your location to the world without realizing it. I've written about this before, e.g., in this post that mentions cybercasing.
It is within the reach of technology to build a "worry-free social sharing" client. The problem is getting the research funding to do so. Industry doesn't really have an interest in having users start being concerned about the implications of their sharing behavior. (It's in their interest to just send the message "share more".) Sure, it's unpleasant and possibly off-putting to have to reflect on the fact that someone might break into your house based on information about your location gleaned from videos that you post to YouTube. But is seems to me that "worry-free sharing" is an idea that users could identify with: just like the cereal box in the morning that announces how much fiber and how many vitamins we are consuming promotes consumption rather than driving people away from a product.
Another project proposal won the competition over the "worry-free social sharing" idea. One of the professors involved in the review later informed me that "worry-free social sharing" sounded like something female. I wasn't really sure what to do with that remark beyond thinking that it probably wasn't one of the considerations for the decision and storing it away for future reference.
I hadn't thought about the femaleness of privacy protection until this weekend, during the new Batman Movie. Here, we watched Cat Woman chasing something called "Clean Slate". She knows that what she needs in order to live her life the way she wants it is to make a clean break with the past. But Batman eventually recognizes this too. And I am happy to see other voices online interested in the privacy themes of the Batman movie. So I am not going to assume that there is only one half of the world population that would be interested in "worry free" sharing solutions.
Thinking about Batman also brought me back to the parallel with the cigarette warning label case. The label pictured above warns of the dangers of second hand smoke, "You're not the only one smoking this cigarette." If warning people about the dangers for their near and dear ones motivates people to cut back or stop smoking, maybe the same effect is true of social sharing. The "worry-free social sharing" client can remind us: Hey, you don't mind posting this picture, but maybe it will have unintended consequences for your friend, who is also pictured.
If you don't believe me, believe Batman: "You wear the mask to protect those you love."
Gerald Friedland, Oriol Vinyals, and Trevor Darrell. Multimodal location estimation. In Proceedings of the international conference on Multimedia (MM '10). ACM, New York, NY, USA, 1245-1252.
Wednesday, July 11, 2012
Time Machine Session at ICME 2012 and beyond
Today was the day of the Time Machine Session at ICME 2012. The session consisted of talks given by experts in the field of multimedia about "Time Machine Topics", defined as: ideas that were published before their time and have yet to reach their full potential.
At first, it might sound like just digging around in the past and brushing off some old ideas. Or it even might sound like some futuristic science recycling scheme, designed to make the most of a limited resource.
But a Time Machine Topic is far from dusty, outdated or rare. Instead, a Time Machine Topic is a topic that is currently experiencing renewed relevance because of subsequent developments in technology and also in our expectations and needs as users.
We think that there are a large number of Time Machine Topics and that some of them bear repeated mention to support the integration of new researchers into the research community and also cross-pollination between related research domains.
The Time Machine Session was born at ICME 2012 because Mercan Topkara and I were appointed under the title "Innovation and Demo Chairs". To be honest, I had never heard of a position called "Innovation and Demo Chair" before. The "Demo" part seemed pretty straightforward, but "Innovation"? What could we possibly offer?
We decided that our innovation should create something for the multimedia community that was new and that served a pressing need. With the Time Machine Session we set out to achieve a number of goals:
We decided that our innovation should create something for the multimedia community that was new and that served a pressing need. With the Time Machine Session we set out to achieve a number of goals:
- Stimulate observation and discussion among researchers.
- Emphasize the benefits of knowing the literature.
- Streamline innovation by reducing redundancy.
- Encourage reproducing and reproducible research.
- Maintain the breadth of the solution space to stimulate new algorithms and approaches
For me personally, a major reason for proposing the Time Machine Session is to create a forum where we publicly and, perhaps a bit ritualistically, demonstrate that we as researchers value knowing the literature and knowing where we have been.
Google Scholar reminds us that we "Stand on the shoulders of giants" and the Time Machine Session gives us as scientists an opportunity to remind ourselves of exactly whose shoulders those are (and there are lots of them).
If Time Machine Sessions exist at conferences (and we hope that there will be more in the future at ICME and elsewhere) we think it will incentivize us as researchers to really study and understand the literature. It will ensure that the "Related Work" sections of our papers are a truly integral part of our research that contributes to the forward movement of our field.
I am making the slides I used for the opening of the Time Machine Session available in the hope that they might be useful for other people who want to hold other Time Machine Sessions elsewhere. In the slides, I discuss the session goals in a bit more detail and use plain language and some great mood-setting images.
I wanted to explicitly point out that the images really made the introduction special, and here I owe much thanks on Auntie K on Flickr, who is so thoughtful to make some of her work available under a Creative Commons license.
The four talks in the ICME 2012 Time Machine Session were the following:
- Dynamic Time Warping's New Youth (Xavier Anguera, Telefonica, Spain )
- Designing Calm Technology (John N.A. Brown, Alpen-Adria Universität Klagenfurt, Austria & Universitat Politècnica de Catalunya, Spain )
- Affective multimedia analysis (Mohammad Soleymani, Imperial College London, UK)
- High Order Entropy Coding, (Wenjun Zeng, University of Missouri, USA)
More information on the talks can be found at the ICME 2012 website's expert talks page.
Also, John N.A. Brown creative a short documentary video at ICME 2013 about the Time Machine Session. The video contains people's reactions to the session and a bit more information on how and why we organized it.
The talks in the Time Machine Session were recorded by videolectures.net and is available at the bottom of the page at http://videolectures.net/icme2012_melbourne/
The opening is here:
Time Machine Session: Introduction
Martha Larson
Also, John N.A. Brown creative a short documentary video at ICME 2013 about the Time Machine Session. The video contains people's reactions to the session and a bit more information on how and why we organized it.
The talks in the Time Machine Session were recorded by videolectures.net and is available at the bottom of the page at http://videolectures.net/icme2012_melbourne/
The opening is here:
Time Machine Session: Introduction
Martha Larson
The original call for proposals for expert talks is repeated below, or read it at: http://www.icme2012.org/CallForPapers_ExpertTalk.php
Time Machine Session
Expert Talks on Innovating the Future Leveraging the Past
IEEE International Conference on Multimedia & Expo (ICME) 2012
11 July, 2012, Melbourne, Australia
Multimedia research is moving ahead in leaps and bounds. In order to pursue the most innovative and productive paths forward, we need an in-depth understanding of where we have already been. The Time Machine Session at the ICME 2012 is dedicated to the principle of improving the future by leveraging valuable insights from the past. The session will consist of a series of expert talks that re-introduce ideas that were published "before their time" and, as a result, were never fully exploited. A "Time Machine Topic" is distinguished by the fact that subsequent technological and social developments have led to a renewal of its relevance, making it currently of critical interest and value to the multimedia research community. A Time Machine Talk covers not only the original idea, but also explains why it currently deserves renewed attention and how it can influence the future of multimedia research. We invite the submission of proposals for oral presentations in the ICME 2012 Time Machine Session.
Time Machine Talks should reflect expert-level understanding of the technological and social developments that have taken place in the field of multimedia and have brought about renewed relevance of past concepts. These developments include, but are not limited to:
- Expansion in the volume, diversity and sources of multimedia content
- Increase in the size, speed and sophistication of distribution networks
- Improvement of computing infrastructures in terms of processing, storage and distribution,
- Growth of the variety and capacity of user end devices
- Development of user expectations for new multimedia applications
In sum, the goals of the Time Machine Session are to stimulate the creative thinking of today's multimedia researchers and to maintain the breadth of the solution space in which we develop new algorithms and approaches. Additionally, we believe that Time Machine Talks can help streamline and defragment the innovation process, by encouraging reproduction and reducing redundancy. Finally, we hope that the Time Machine Session will stimulate interesting and productive discussion in the community.
Selection
From the pool of submissions, a panel will make a selection of talks for presentation at the Time Machine Session. The decision will be made on using the following criteria:
From the pool of submissions, a panel will make a selection of talks for presentation at the Time Machine Session. The decision will be made on using the following criteria:
- Renewed relevance of the idea for today's multimedia researchers and research domains as set out in the general ICME 2012 CFP
- Scope of the potential impact of the re-introduction of the idea on innovation in the multimedia research community
- Importance of re-introduction of the idea to prevent the community from wasting time by "reinventing the wheel"
- Presentation of the idea i.e., compelling argumentation and engaging presentation style
Submission Format
The submission consists of three parts:
- The reference (i.e., bibliographic citation) of the paper that originally introduced the idea (pub-lished at least five years ago and still publicly available),
- A three minute video summarizing the idea and explaining why at the present moment its time is finally ripe,
- A 300-400 word abstract to accompany the Time Machine Talk in the ICME 2012 program. Note that the person submitting the proposal does not necessarily need to be one of the authors of the original paper.
Friday, June 1, 2012
Criteria for judging a demo in a conference demo session
Mercan Topkara and I are the "Innovation and Demo Chairs" for ICME 2012 to be held 9-13 July 2012 in Melbourne, Australia. We were called upon to organize the decision making process by which the ICME organizers would arrive at the decision of which demo would take home the ICME 2012 best demo award.
The decision is a difficult one because demos in the area of multimedia tend to be radically different in nature. For this reason, I formulated a list of six dimensions to use when judging demos.
1. Clarity: Understandability of the demo paper and the presentation.
2. Realization: Well implemented, robust, good use of technology.
3. Innovation: Addresses a problem that has not yet been tackled (or has proven difficult to solve).
4. Impact: The number of people the technology potentially touches and the importance of its influence on their lives.
5. Representativity: Centrality to the topics covered by the conference (in this case ICME)
6. Magic: How closely the technology fills the description, "Any sufficiently advanced technology is indistinguishable from magic" (cf. http://en.wikipedia.org/wiki/Clarke%27s_three_laws)
Number 6 is basically a wild card that makes it possible to introduce in a controlled way that factor of je ne sais quoi, which seems to slip into the considerations made when judging demos in any case.
In practice, another factor that always seems to be important is how close the demo is to a working system that is or is about to be deployed in the real world. Also, when judging demos it seems that one is always trying to project forward: how important will this technology be five or ten years from now? Will the passage of time reveal that it is a disruptive technology? (Or, as Wikipedia prefers to call it disruptive innovation?)
For a list of the demos to be presented at ICME 2012, see the ICME 2012 Demo page.
I am dating the post 1 June, when I formulated the list of criteria. It's later now, but time seems to have simply gotten away from me, not surprising given the ICME 2012 Time Machine Session.
The decision is a difficult one because demos in the area of multimedia tend to be radically different in nature. For this reason, I formulated a list of six dimensions to use when judging demos.
1. Clarity: Understandability of the demo paper and the presentation.
2. Realization: Well implemented, robust, good use of technology.
3. Innovation: Addresses a problem that has not yet been tackled (or has proven difficult to solve).
4. Impact: The number of people the technology potentially touches and the importance of its influence on their lives.
5. Representativity: Centrality to the topics covered by the conference (in this case ICME)
6. Magic: How closely the technology fills the description, "Any sufficiently advanced technology is indistinguishable from magic" (cf. http://en.wikipedia.org/wiki/Clarke%27s_three_laws)
Number 6 is basically a wild card that makes it possible to introduce in a controlled way that factor of je ne sais quoi, which seems to slip into the considerations made when judging demos in any case.
In practice, another factor that always seems to be important is how close the demo is to a working system that is or is about to be deployed in the real world. Also, when judging demos it seems that one is always trying to project forward: how important will this technology be five or ten years from now? Will the passage of time reveal that it is a disruptive technology? (Or, as Wikipedia prefers to call it disruptive innovation?)
For a list of the demos to be presented at ICME 2012, see the ICME 2012 Demo page.
I am dating the post 1 June, when I formulated the list of criteria. It's later now, but time seems to have simply gotten away from me, not surprising given the ICME 2012 Time Machine Session.
Thursday, May 17, 2012
Search by misconception: Should search engines support information needs that are ill conceived?
![]() |
| The real Mozartkugeln? (Flickr: davidroethler) |
Let's start with an example. When I refer to "Mozartkugeln" I mean the ones in the pictured here. They are gold and show Mozart in his red jacket. Their shape is round. I differentiate these from the ones with flat bottoms, which are for me "fake" Mozartkugeln.
My idea of Mozartkugeln can be considered a misconception. The original Mozartkugeln are apparently produced by a company called "Fürst" and are silver with a blue Mozart. Additionally, the producer of the flat-bottom ones apparently has the right to call their product "Real Reber Mozartkugeln". Digging on Wikipedia and on other websites supplied me with this information.
But what should an image search engine return in response to the query "Mozartkugeln"? Is it obliged to make an effort to resolve the question of which is the "real" Mozartkugel? Or is it fine if it just returns images that users have uploaded a tagged with "Mozartkugel"?
Effectively, simply returning images tagged with "Mozartkugel" allows users to search by misconception. The search engine returns images who have been tagged by people like me, who have a certain view on the matter (based on conversations with an Austrian roommate now a couple decades old and several subsequent trips to Austria, none including Salzburg), which is not necessarily universal. I am not immediately convinced that I can be satisfied with such a search engine as a source of information. Although, it seems reasonable to assume that if enough voices are combined, a consensus will emerge. I noticed that if you search for "champagne" on Google images, the top hits (at least the ones that depict identifiable bottles) clearly hail from the Champagne region in France and don't include the large range of other bubbly wines from other corners of the world that are widely enjoyed under the name "champagne".
In short, allowing search by misconception seems relatively innocuous. But we should be careful about assuming that the Mozartkugeln example is the end of the story. What is unique about this example, is that the search engine is relatively transparent in the way that it works. The images are returned by seeking exact matches in their user-assigned tagsets; without such a match, the image is not relevant. Users of the search engine have a chance of being at least vaguely aware of the reason for the match and they can propagate their understanding of the reliability of the taggers to create an understanding about the reliability of the results.
However, when the search engine becomes more sophisticated, the situation quickly gets quite murky. For example, if I had a visual concept detector that was trained to detect Mozartkugeln in images and assign to them the appropriate tags. The design of the detector would require collecting examples of Mozartkugeln, which means that whoever trains the detector holds the ultimate control over deciding what a Mozartkugel actually is.
The example of Mozartkugeln is interesting. In some cases, one could argue that common sense knowledge will tell you what an object is, for example, a helicopter or a pram. Everyone can identify these objects, right? But in the case of the Mozartkugeln, there is no right answer. It depends on your perspective. A long discussion will arrive at the conclusion "It's complicated". (And you may already find yourself with the same issue for the pram, if not actually the helicopter.)
It seems like a good idea to do away with the central authority that collects the examples used to train the detectors. After all, no one really likes they guy who walks around the party reminding people, "Yes, but it's not real champagne".
But do we really want to admit search by misconception? I had quite an unsettling experience with Google's query suggestion. On 17 May 2011, I was looking for a news story on one of the Facebook founders have renounced his US citizenship. No sooner had I typed in "facebook founder" did Google present me with the following list of suggested queries:
facebook founder mark
facebook founder saverin
facebook founder buys new republic
facebook founder college
facebook founder gay
facebook founder bios
facebook founder dead
facebook founder movie
Did I really need to know about the existence of the circulating rumors? Do I go on to passively "believe" a query or do I dig deeper that find out if it is true? Do we really want our search engines to allow us to so easily flow down
the same information paths worn by searchers before us who mis-received a
rumor?
In the case of "facebook founder dead" I did dig deeper. That query led to a Fox News article on the death of Ilya Zhitomirskiy, one of the co-founders of Diaspora*, an alternative to Facebook. I was left wondering at how query suggestions have taken on an information dissemination (news broadcast, if you will) role of their own.
In the case of "facebook founder dead" I did dig deeper. That query led to a Fox News article on the death of Ilya Zhitomirskiy, one of the co-founders of Diaspora*, an alternative to Facebook. I was left wondering at how query suggestions have taken on an information dissemination (news broadcast, if you will) role of their own.
From the fun of searching for pictures of bonbons (...and wondering if round vs. flatbottom Mozartkugel relates to a real misconception or rather an alternate interpretation) we hit on a matter of true importance (Diaspora* upends the Facebook model because it is based on the idea that every member of the network should “own” his personal information). Suddenly, it gets extremely serious. In light of this seriousness, it looks like if we really do not want search engines to admit search by misconception at all.
This whole line of though was started while I was at a symposium entitled "Cultural Heritage Gets Social" of the SEALINCMedia project (Socially-enriched access to linked cultural media). Alice Warley (Public Catalogue Foundation, UK) gave a talk entitled "Your Paintings Tagger: Crowd-sourcing, art history and the UK's national oil painting collection" about a website where visitors collaborate to tag painters.
Apparently, general public users tend to tag older paintings "formal wear" when what the people pictured in the paintings are wearing is not formal wear at all, but rather daily clothing.
The reason for this misconception is that today's formal wear evolved from what was worn on a daily basis in certain social circles in past eras. Wikipedia is rather silent on the history of formal wear. The "misconceptions" of the taggers are actually a source of information about something that is not widely known, but actually an real historical connection.
The reason for this misconception is that today's formal wear evolved from what was worn on a daily basis in certain social circles in past eras. Wikipedia is rather silent on the history of formal wear. The "misconceptions" of the taggers are actually a source of information about something that is not widely known, but actually an real historical connection.
So we're back to the Mozartkugeln, considering whether Mozart is dressed for a concert, or is in his everyday work clothing that he uses for composing. It seems like misconceptions help us to uncover new and interesting information.
However, if we incorporate misconceptions, maybe we should call them 'exploration engines'. A 'search' engine should find answers or else gently reveal to us that our initial information need was ill conceived.
However, if we incorporate misconceptions, maybe we should call them 'exploration engines'. A 'search' engine should find answers or else gently reveal to us that our initial information need was ill conceived.
Subscribe to:
Posts (Atom)



