Friday, February 15, 2013

Crowdsourcing for Multimedia: ACM Multimedia and The Crowd

Crowdsourcing for multimedia is a set of techniques that leverage human intelligence and a large number of individual contributors in order to tackle challenges in multimedia that are conventionally approached using automatic methods. Exploiting the crowd means taking advantage of human computation where it can help support multimedia algorithms and multimedia systems the most.

ACM Multimedia 2013 has introduced a new Crowdsourcing for Multimedia area:
http://acmmm13.org/submissions/call-for-papers/#crowdsourcing
The area cuts clean across traditional multimedia areas, touching upon nearly every topic relevant for multimedia. The area casts a wide net to include the full range of research results and novel ideas in multimedia that are made possible by the crowd, i.e., they exploit crowdsourcing principles and techniques.

How did this new area arise? Crowdsourcing's grand debut at ACM Multimedia is CrowdMM http://www.crowdmm.org/. On Octber 29, 2012 the CrowdMM 2012 International ACM Workshop on Crowdsourcing for Multimedia was held in conjunction with ACM Multimedia 2012 in Nara, Japan. The workshop kicked-off with a keynote entitled "PodCastle and Songle: Crowdsourcing-Based Web Services for Spoken Content Retrieval and Active Music Listening" by Masataka Goto of the National Institute of Advanced Industrial Science and Technology (AIST), Japan. These two systems dazzled the audience and gave us a foretaste of the possibilities that the power of the crowd opens for the multimedia community. An interesting day of talks, posters and discussion ensured, culminating in a panel (summarized below).

The organizers of CrowdMM 2012 hope that both the ACM Multimedia area (focused on groundbreaking research results) and also CrowdMM workshop (focused on methodology, exploratory work and on researcher interaction) will provide a solid foundation that allows crowdsourcing for multimedia to grow within the multimedia community to reach its full potential.

"Crowdsourcing for multimedia: At a crossroad or on a superhighway?"

Summary of the Panel Discussion at CrowdMM 2012 

What is the potential of crowdsourcing for ACM Multimedia?
We need The Crowd to allow us to build larger, up-to-date dictionaries for multimedia annotation. We also need The Crowd to create ground truth at a large scale.

The combining techniques for active learning and for incentivizing human contributions will contribute to many different specific multimedia problems.

In all cases, both quality control will be important and also making it fun for The Crowd to contribute, e.g., continuing to build entertaining games to collect Crowd contributions. User engagement breeds quality: for example, Songle provides services that are enjoyable to use and attracts good workers naturally. 
 
What are the limitations of crowdsourcing for multimedia?
Data from non-experts is valuable, but for some tasks we need experts. We need methods that will allow us to identify experts, for example, with domain knowledge. The multimedia community can potentially address the problem of having access to "the right crowd", by joining forces to cultivate a community of crowdsourcing workers who deliver high quality annotations for specific multimedia domains.

How would ACM Multimedia be different had crowdsourcing been invented 20 years ago? If crowdsourcing had existed 20 years ago, we would make much more effective use of active learning paradigms, i.e., algorithms that would interactively query human annotators to obtain new labels for certain multimedia items.

Crowdsourcing makes possible large scale multimedia annotations. Even if crowdsourcing existed 20 years ago, we may not have the tools and techniques to deal with large scale data.

The challenge today is to realize the potential of multimedia, both in venturing into new domains for research and also in scaling up our systems to exploit larger amounts of human labeled data for training and also for evaluation.

In short, the panel concluded, it’s up to ACMMM to catch up with the crowd.

A big thank you to my fellow organizers for the work that they put into making CrowdMM 2012 such a success:
IMG_0333

Wednesday, January 2, 2013

Brave New Tasks: Incubating Innovation in the MediaEval Multimedia Benchmark

An important innovation of MediaEval 2012 was the "Brave New Task" track. MediaEval is a multimedia benchmark that offers and promotes challenging new tasks in the area of multimedia access and retrieval. We focus on tasks that emphasize multiple modalities ("The 'multi' in multimedia") and that have social and human relevance.

Brave New Tasks were introduced because we noticed that there is rather a tension between benchmark evaluation and innovation. Benchmarking is essentially a conservative activity: we want to compare algorithms on the same task using the same data set and the same evaluation procedure. This sameness allows us to chart progress with respect to the state of the art, especially over the course of time. How do we innovate, when the key strength of benchmarking is that we repeatedly do the same thing?

We innovate by tackling new problems. However, in order to create a successful benchmarking task from a new problem, a number of questions must be answered. Is the problem suitable for evaluation in a benchmark setting? What sort of data is needed to evaluate solutions developed by benchmark participants? How much effort is needed to create ground truth? Do we need to refine our definition of the task and of the evaluation procedure? Is there an actual chance that algorithms can be developed to solve the task and what resources are needed? Is there a critical mass of interest in solving this problem? Are the solutions appropriate for application?

The easy way forward would be to insist that there are clear answers to all of these questions prior to running a task. In some cases, is will be possible to gather the answers. In others, however, it will not. Forcing tasks to have answers before attempting to create a benchmark poses a serious risk that researchers will avoid the truly challenging and innovative tasks because they receive the message that they need to "play it safe."

People in the MediaEval community rattle their swords and shields when they are told that they need to "play it safe." Brave New Tasks support innovation in MediaEval by incubating tasks in their first year, allowing the task organizers to answer these questions. We value the advantages that the conservative aspects of benchmarking bring to the community, but we also thrive by taking risks. The Brave New Task track is a lightly protected space that allows us to take the risks that allow our benchmark to continue to renew itself.

Because people have asked me about Brave New Tasks in MediaEval 2013 and "How did you do it?" I am providing here a more detailed description of how it works and how we anticipate that it will develop in 2013:

To start, let me write a few words about makes a main "mainstream" task in MediaEval (i.e., a task that is not a Brave New Task). At the end of the calendar year, MediaEval solicits proposals from teams who are interested in organizing a task in the next MediaEval season. For an example, see the MediaEval 2013 call for task proposals. Whether a proposal is accepted as a MediaEval task depends on the interest expressed on the MediaEval survey. The survey is published in the first days of January and circulated widely to the larger research community.

During the survey, task proposers gather information on who is interested in carrying out their tasks. By the time the survey concludes, the proposers must have promises from five core participants (who are not themselves organizers) who will cross the finish line of the task (including submitting, results writing the working notes paper and attending the MediaEval workshop) come "hell or high water". This selection criteria is set up so that we have a minimum number of results to compare across sites for any given task---if there are only one or two, we don't get the "benchmark effect".

Tasks that the community finds interesting and promising, but that do not necessarily meet these stringent selection criteria, can be selected as Brave New Tasks. The difference between a Brave New Task and other MediaEval tasks is that these tasks are new, and ideally also a scientifically risky (in the responsible sense of "risky").

Brave New Tasks are run "by invitation only". The "invitation only" clause does not make the task exclusive: anyone who asks the task organizers can be granted an invitation. Instead, the clause allows the tasks to handle unexpected situation by, if necessary, decoupling their schedules from the main task schedules to accommodate unforeseen delays in data set releases. Participants of past editions of MediaEval will recognize the usefulness of a mechanism that makes the benchmark robust to unexpected situations.

Further,  Brave New Tasks do not require their participants to submit working notes papers or attend the workshop. The "only" requirement that the task must fulfill is to contribute an overview paper in the MediaEval working notes proceedings that sums up the task and presents and outlook for future years. One or more of the organizers attends the workshop to make the presentation and participate in the discussion about whether the task should target developing into a mainstream task in the next year.

A "Brave New Task" is encouraged to go far beyond the minimum requirement. In fact, 2012 saw one of the Brave New Tasks "Search and Hyperlinking" achieve the scope of a mainstream task, with six working notes papers from task participants appearing in the MediaEval 2012 working notes proceedings. The task was effectively indistinguishable from mainstream tasks in its contribution to the benchmark.

In 2013, we plan to strengthen the Brave New Task track by providing them with more central support. The tasks will be run under the same infrastructure as the mainstream tasks and decoupled from the schedule only if it is absolutely necessary. They will also be given the option of using the central registration system.

Brave New Tasks have been a successful innovation in MediaEval 2012 and one that we hope to strengthen in the future. I'd like to end by pointing out that it is not so much the "rules" of Brave New Tasks that have made them such a success, but rather the efforts of the Brave New Task organizers. Success is dependent on having a group of devoted researchers with a vision for a new task idea and the capacity and stamina to see it through the first year...including long hours spent reading related work, developing new evaluation metrics (if necessary), contacting and following up with participants, collecting data and creating ground truth. It is not so much the tasks themselves that are brave, but the organizers who are fearless and relentless in their pursuit of innovation.

Forward, charge!

Friday, November 30, 2012

ImageNet and the Edge of the World: On visual concept labels for images


Google Image Search results for the query "two-year-old horse"
ImageNet (http://www.image-net.org/) is a collection of images depicting concepts in the lexical database WordNet (http://wordnet.princeton.edu/). ImageNet consists of groups of images that illustrate the same WordNet concept. On WordNet a concept is a set of cognitive synonyms, or words that are understood to express the same thing.

ImageNet is very cool. If I were a kid, this would have been better than any of those other picture dictionaries...it is fun just to click through and explore what exists in the world.

I recently made a video about this paper:

Jia Deng; Wei Dong; Socher, R.; Li-Jia Li; Kai Li; Li Fei-Fei; , "ImageNet: A large-scale hierarchical image database," Proc. IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2009.  pp.248-255.

The video's down below, in case you want to get a quick overview of the paper...but in the video I am mostly focusing on discussing the crowdsourcing methods applied to create ImageNet.

Here, I will discuss "Edge of the ImageNet Image World". I identify the edge with the idea of a concept being "difficult to be illustrated". This concept I found mentioned in footnote 1 of the  paper:

About 20% of the synsets have very few images, because either there are very few web images available, e.g. “vespertilian bat”, or the synset by definition is difficult to be illustrated by images, e.g. "two-year-old horse". (p. 249)

I would like to make the point that we may be radically underestimating the importance of "difficult to be illustrated" visual concepts in our multimedia information indexing systems.

Difficult to be illustrated? It is quite obvious that it is not inherently difficult to have a picture of a two-year-old horse.

I have a horse, two years ago I watched it being born, and I take a picture of it. Finished.

What is difficult is to find a group of people (annotators or users) who will look only at the picture (knowing nothing about me and the horse) and agree that the horse is two years old.

It is difficult for two reasons:
  1. Context of use: The concept "two-year old horse" is difficult to pin down exactly. Does a horse that is two-years and one day old still count as a two-year-old horse? It depends on what you are using the picture for. If you are using it for a collection of "horses on their second birthdays" it won't count. However, if you are using it to illustrate horses that are less than full grown, that day doesn't matter.
  2. Background of user: You have to know something about horses to distinguish a horse that is a foal (under one year) from one that is a colt or a filly (which Wikipedia tells us are terms that may be used until the horse is 3 or 4).
The ImageNet paper claims that "ImageNet aims to provide the most comprehensive and diverse coverage of the image world".

As multimedia researchers, we seem to assume that these "difficult to illustrate concepts" represent some marginal part of multimedia meaning. I mean, it's less that 20% of the concepts in WordNet that have this problem, so isn't it a good first approximation to just ignore them and focus on the 80% that are easily illustrated by images?

Context of use: We can just concentrate on the formal definitions of concepts. It's about delivering precise results lists when we search for images isn't it? Under that view, we can solve point 1. by deciding to use the most restrictive definition possible: the horse that turned three yesterday is no longer a two-year-old horse.

OK. So we all are totally annoyed at the guy who just sits immobile when we say, "Hand me that red screwdriver?" You climb down from the ladder just to hear him say, "I see a crimson screwdriver, but no red screwdriver." We are annoyed because we know that language is built to be used, and part of that use is the fact that we accommodate the meanings of words within their contexts of use.

But we learn to live with it. We realize the guy is literally right, so we grab the screwdriver ourselves and climb back up the ladder. We could learn to live with image search engines that behave like that as well, couldn't we?

Background of user: We can just concentrate on what the "man on the street" thinks about the image. It's about delivering results that are generally recognizable and not results that require some expert insight, isn't it? Under that view we can solve point 2. by deciding to use what a member of the general public would say about the image: it's a horse, probably not a grown up horse, but there's no telling if it's two years old.

Whoa. Hold your horses right there! Who gets to then decide who constitutes the "man on the street" of the "general public"?

Many people that I meet on the street in my daily life are not going to know the difference between quite obvious concepts like "bananas" vs. "plantains". It depends on what street I chose to look at.

With respect to many streets in Western Europe, "plantain" would be "difficult to be illustrated": people that can identify them are somehow considered experts. Not so in West Africa.

Irresponsible intuitions: In a split second, we as multimedia researchers can make a decision that seems "obvious", but that on closer consideration has potential to come back and haunt us.

We are reinforced to make these "obvious" decisions because they are the ones that allow us to continue on with our research with a minimal investment of resources in creating labeled image sets.

If I use restrictive, formal decisions, I don't have to turn to actual users of image search engines to try to understand how the "language of concepts" that they use when they search.

I also don't have to try to dig down to more subtle forms of cultural bias that exist in WordNet. Who of us has time to read a volume on cultural bias in dictionaries with contributions from 40 scholars?

In the end, although "difficult to be illustrated concepts" may constitute 20% of the concepts in WordNet, we have no idea of what percent of actual user image need might be related to these concepts. It could be huge!

Edge of the world: Google somehow gets it right. The search results at the top are returned by Google Images in response to the query "two-year-old horse". The first image occurs on the Internet in conjunction with the text "2 Year old Buckskin Quarter horse Colt". Someone apparently took a picture of their two-year-old horse and that seems to be right.

In the next picture, it's the kid and not the horse that's two, but that's pretty obviously wrong, and even amusing.

At the very least, this discussion allows us that to conclude that if ImageNet covers "The Image World", that is a very flat world indeed. It is easy to follow a "difficult to be illustrated" concept to the end of that world and stand there looking over the edge...

...ImageNet is a valuable research tool and serves the community well. However, we should all be aware of exactly where the edge of the ImageNet world is, not that we want to avoid it, but perhaps because that is exactly the place from which we want to leap off.


Saturday, August 25, 2012

Gender in Advertising Images: The Devil is in the Detail

My Saturday was unbalanced already at breakfast, while reading the Economist and drinking my orange juice. On page 67 of the August 25-31, 2012 issue, I discovered that TU Delft is recruiting a Professor of Safety Science (good news!). Unfortunately, whoever designed this ad has made some unconventional decisions (not such good news).

The most obvious "bug" is the choice to include in the advertisement an image of a person. Since this misstep is a useful illustration of the limitations of visual depictions in multimedia, I decided to dedicate a blogpost to discussing it.

At first consideration, it seems obvious that our university should advertise using image of people. One of the reasons that I love working at TU Delft is the emphasis on solving societal problems. Using pictures containing people and not just technology wherever possible seems to be a good strategy for getting the importance of our work to address human and social challenges across.

However, a major limitation for visual depictions such as images and videos is "the curse of instance  depiction". Basically, it is impossible to create such visual imagery without committing yourself to depicting a full range of details. You can't get across and abstract concept, for example, "car" without actually committing yourself to an instance of a single car existing in the real world, which you take to stand for all cars. Instead, you are going to need to show in your image a specific type, make and model.

Here, the concept that the ad is trying to convey is "professor". The "type, make and model" chosen to convey this concept are an adult of a certain gender and a certain age group, wearing glasses. It seems plausible that the person designing the ad was aware of the problem of instance depiction. The decision to use a model with a shaved head makes it possible to avoid depicting the hair color, which could serve to further specify the ethnic background or the age.

However, it is extremely difficult, if not impossible, to "hedge" on the gender question in images of the real world. A person depicted in a daytime work setting will generally be identifiable as a male person or a female person.

If we assume that the process by which we choose and interpret images that are being used to represent categories follows prototype theory, then the choice of a male to represent a TU Delft professor is no just unbalancing for the reader of the advertisement, but is very serious indeed. Prototype theory tells us that in our cognitive representations, some members of conceptual categories are more salient than others. We think of them first when we think of a category and we react to them more quickly when confirming category membership.

The use of a male person in this advertisement sends the message that males are the canonical professors at the TU Delft. Although men are clearly in the majority in the faculty, there is not any sort of a conscious intention at the university to keep the situation that way. In fact, I have the impression that everyone is working to shift their idea of how can be a professor to encompass a diverse demographic more directly representative of the general population.

Visual depictions in multimedia, i.e., images depicting the real world, are limited in what they can express because they deprive us of the possibilities of leaving certain details unpecified. What we have is a reversal of the saying "A picture is worth a thousand words." Instead, the spoken or the written word is able to express more in this case because human language can directly convey concepts without having to make use of specific instances to do so. In effect, the possibility for ambiguity or underspecification is makes human language more expressive that multimedia.

And so, the saying "The devil is in the detail" takes on a new shade of meaning.

What to do about the advertisement? I advise having a closer look at some advertising guidelines. Advertising Standards Canada, a non-profit self-regulation body for advertising, has a helpful list of guidelines for balancing gender representations in advertising online and surely Europe has a similar set of guidelines.

An "quick and dirty" solution is to look to see how other universities advertise. In the Economist, a general tendency to avoid imagery is readily apparent. For example, next to the TU Delft advertisement is a classical advertisement for Harvard faculty positions, whose only graphic content in the Harvard Business School logo.

I was cheered up again when my Google Googles app confirmed for me that the logo used was from the business school (i.e., distinct from the main Harvard Logo). It is my first use of Google Goggles for something other than just playing around with while hanging out with my multimedia information retrieval colleagues.

For completeness, I note a less obvious bug. The advertisement contains the text "Maximum employment: 38 hours per week (1 FTE)" In order to interpret this text, you need to know that "FTE" stands for "full time equivalent". 1 FTE means this position is a full time job. Contrary to what the text implies, no one the Safety Science processor to working 38 hours a week.

Wednesday, August 8, 2012

Worry-Free Social Sharing for Social Networks

 Flickr: Phil Wiffen
Quite a few people have heard me say that social networks should come with a warning: when you sign up, for example for Facebook, the company should be required to notify you of the danger of long term impact of social  sharing on your personal privacy (and sometimes I also add two other factors: how many hours you are projected to spend "Facebooking" over the coming years and also how much peer pressure and social isolation you will endure if you want to leave the platform). In my lifetime, awareness has developed and legislation has changed such that cigarettes and cigarette advertisements are required to bear warnings about the health hazards: maybe I'll yet live to see awareness rise about the consequences of social sharing.

In contrast to smoking, social sharing done right actually helps rather than hurts. In fact, the rise of online social networking and social multimedia sharing has been downright amazing technological development. Moved by this awe, last year in a project proposal, I effused that social networks are, "...a virtual prosthetic that extends the strong fabric of social connectivity critical to the well-being and growth of human societies into the online realm."

That proposal developed the idea of "worry-free social sharing": a social sharing client that would gently alert us when our sharing actions, in ways we do not intend, threaten to compromise our privacy---and then suggest alternative actions, which allow us to share our personal experiences, but in a wiser way.

Yes, people are responsible for their own actions. But in some cases, we as individual users do not have the understanding of multimedia analysis technology, or of the power of algorithms to combine different sorts of data to reveal facts about us that we thought were hidden. We all would need such understanding in order to allow us to make informed decisions about which types of social sharing is harmless and which types should better be avoided.

Even for the most savvy of us there are always surprises: Did you know that if you upload a video to YouTube and you carefully avoid geo-tagging it, but if you happen to be in a city and capture an ambulance siren in the background, that siren will serve to indicate in which city you are? Check out the work on multimodal location estimation [1]. Maybe you don't care if the world knows where you are, but if you do happen to be worried about having left your house empty during your vacation, it would be good to know that you just about betrayed your location to the world without realizing it.  I've written about this before, e.g., in this post that mentions cybercasing.

It is within the reach of technology to build a "worry-free social sharing" client. The problem is getting the research funding to do so. Industry doesn't really have an interest in having users start being concerned about the implications of their sharing behavior. (It's in their interest to just send the message "share more".) Sure, it's unpleasant and possibly off-putting to have to reflect on the fact that someone might break into your house based on information about your location gleaned from videos that you post to YouTube. But is seems to me that "worry-free sharing" is an idea that users could identify with: just like the cereal box in the morning that announces how much fiber and how many vitamins we are consuming promotes consumption rather than driving people away from a product.

Another project proposal won the competition over the "worry-free social sharing" idea. One of the professors involved in the review later informed me that "worry-free social sharing" sounded like something female. I wasn't really sure what to do with that remark beyond thinking that it probably wasn't one of the considerations for the decision and storing it away for future reference.

I hadn't thought about the femaleness of privacy protection until this weekend, during the new Batman Movie. Here, we watched Cat Woman chasing something called "Clean Slate". She knows that what she needs in order to live her life the way she wants it is to make a clean break with the past. But Batman eventually recognizes this too. And I am happy to see other voices online interested in the privacy themes of the Batman movie. So I am not going to assume that there is only one half of the world population that would be interested in "worry free" sharing solutions.

Thinking about Batman also brought me back to the parallel with the cigarette warning label case. The label pictured above warns of the dangers of second hand smoke, "You're not the only one smoking this cigarette." If warning people about the dangers for their near and dear ones motivates people to cut back or stop smoking, maybe the same effect is true of social sharing. The "worry-free social sharing" client can remind us: Hey, you don't mind posting this picture, but maybe it will have unintended consequences for your friend, who is also pictured.

If you don't believe me, believe Batman: "You wear the mask to protect those you love."

Gerald Friedland, Oriol Vinyals, and Trevor Darrell. Multimodal location estimation. In Proceedings of the international conference on Multimedia (MM '10). ACM, New York, NY, USA, 1245-1252.

Wednesday, July 11, 2012

Time Machine Session at ICME 2012 and beyond

Today was the day of the Time Machine Session at ICME 2012. The session consisted of talks given by experts in the field of multimedia about "Time Machine Topics", defined as: ideas that were published before their time and have yet to reach their full potential. 

At first, it might sound like just digging around in the past and brushing off some old ideas. Or it even might sound like some futuristic science recycling scheme, designed to make the most of a limited resource.  

But a Time Machine Topic is far from dusty, outdated or rare. Instead, a Time Machine Topic is a topic that is currently experiencing renewed relevance because of subsequent developments in technology and also in our expectations and needs as users. 

We think that there are a large number of Time Machine Topics and that some of them bear repeated mention to support the integration of new researchers into the research community and also cross-pollination between related research domains.

The Time Machine Session was born at ICME 2012 because Mercan Topkara and I were appointed under the title "Innovation and Demo Chairs". To be honest, I had never heard of a position called "Innovation and Demo Chair" before. The "Demo" part seemed pretty straightforward, but "Innovation"? What could we possibly offer? 

We decided that our innovation should create something for the multimedia community that was new and that served a pressing need. With the Time Machine Session we set out to achieve a number of goals:
  1. Stimulate observation and discussion among researchers.
  2. Emphasize the benefits of knowing the literature.
  3. Streamline innovation by reducing redundancy.
  4. Encourage reproducing and reproducible research.
  5. Maintain the breadth of the solution space to stimulate new algorithms and approaches
For me personally, a major reason for proposing the Time Machine Session is to create a forum where we publicly and, perhaps a bit ritualistically, demonstrate that we as researchers value knowing the literature and knowing where we have been. 

Google Scholar reminds us that we "Stand on the shoulders of giants" and the Time Machine Session gives us as scientists an opportunity to remind ourselves of exactly whose shoulders those are (and there are lots of them). 

If Time Machine Sessions exist at conferences (and we hope that there will be more in the future at ICME and elsewhere) we think it will incentivize us as researchers to really study and understand the literature. It will ensure that the "Related Work" sections of our papers are a truly integral part of our research that contributes to the forward movement of our field.

I am making the slides I used for the opening of the Time Machine Session available in the hope that they might be useful for other people who want to hold other Time Machine Sessions elsewhere. In the slides, I discuss the session goals in a bit more detail and use plain language and some great mood-setting images. 

I wanted to explicitly point out that the images really made the introduction special, and here I owe much thanks on Auntie K on Flickr, who is so thoughtful to make some of her work available under a Creative Commons license.

The four talks in the ICME 2012 Time Machine Session were the following:
  • Dynamic Time Warping's New Youth (Xavier Anguera, Telefonica, Spain )
  • Designing Calm Technology (John N.A. Brown, Alpen-Adria Universität Klagenfurt, Austria & Universitat Politècnica de Catalunya, Spain )
  • Affective multimedia analysis (Mohammad Soleymani, Imperial College London, UK)
  • High Order Entropy Coding, (Wenjun Zeng, University of Missouri, USA)
More information on the talks can be found at the ICME 2012 website's expert talks page.

Also, John N.A. Brown creative a short documentary video at ICME 2013 about the Time Machine Session. The video contains people's reactions to the session and a bit more information on how and why we organized it.



The talks in the Time Machine Session were recorded by videolectures.net and is available at the bottom of the page at http://videolectures.net/icme2012_melbourne/ 

The opening is here:
   
Time Machine Session: Introduction

Martha Larson

The original call for proposals for expert talks is repeated below, or read it at: http://www.icme2012.org/CallForPapers_ExpertTalk.php 

Time Machine Session 
Expert Talks  on Innovating the Future Leveraging the Past
IEEE International Conference on Multimedia & Expo (ICME) 2012
11 July, 2012, Melbourne, Australia

Multimedia research is moving ahead in leaps and bounds. In order to pursue the most innovative and productive paths forward, we need an in-depth understanding of where we have already been. The Time Machine Session at the ICME 2012 is dedicated to the principle of improving the future by leveraging valuable insights from the past. The session will consist of a series of expert talks that re-introduce ideas that were published "before their time" and, as a result, were never fully exploited. A "Time Machine Topic" is distinguished by the fact that subsequent technological and social developments have led to a renewal of its relevance, making it currently of critical interest and value to the multimedia research community. A Time Machine Talk covers not only the original idea, but also explains why it currently deserves renewed attention and how it can influence the future of multimedia research.  We invite the submission of proposals for oral presentations in the ICME 2012 Time Machine Session.

Time Machine Talks should reflect expert-level understanding of the technological and social developments that have taken place in the field of multimedia and have brought about renewed relevance of past concepts. These developments include, but are not limited to:
  • Expansion in the volume, diversity and sources of multimedia content
  • Increase in the size, speed and sophistication of distribution networks
  • Improvement of computing infrastructures in terms of processing, storage and distribution,
  • Growth of the variety and capacity of user end devices
  • Development of user expectations for new multimedia applications
In sum, the goals of the Time Machine Session are to stimulate the creative thinking of today's multimedia researchers and to maintain the breadth of the solution space in which we develop new algorithms and approaches. Additionally, we believe that Time Machine Talks can help streamline and defragment the innovation process, by encouraging reproduction and reducing redundancy. Finally, we hope that the Time Machine Session will stimulate interesting and productive discussion in the community.

Selection
From the pool of submissions, a panel will make a selection of talks for presentation at the Time Machine Session. The decision will be made on using the following criteria:
  • Renewed relevance of the idea for today's multimedia researchers and research domains as set out in the general ICME 2012 CFP
  • Scope of the potential impact of the re-introduction of the idea on innovation in the multimedia research community
  • Importance of re-introduction of the idea to prevent the community from wasting time by "reinventing the wheel"
  • Presentation of the idea i.e., compelling argumentation and engaging presentation style 
Submission Format
The submission consists of three parts:
  1. The reference (i.e., bibliographic citation) of the paper that originally introduced the idea (pub-lished at least five years ago and still publicly available),
  2. A three minute video summarizing the idea and explaining why at the present moment its time is finally ripe,
  3. A 300-400 word abstract to accompany the Time Machine Talk in the ICME 2012 program. Note that the person submitting the proposal does not necessarily need to be one of the authors of the original paper.

Friday, June 1, 2012

Criteria for judging a demo in a conference demo session

Mercan Topkara and I are the "Innovation and Demo Chairs" for ICME 2012 to be held 9-13 July 2012 in Melbourne, Australia. We were called upon to organize the decision making process by which the ICME organizers would arrive at the decision of which demo would take home the ICME 2012 best demo award.

The decision is a difficult one because demos in the area of multimedia tend to be radically different in nature. For this reason, I formulated a list of six dimensions to use when judging demos.

1. Clarity: Understandability of the demo paper and the presentation.
2. Realization: Well implemented, robust, good use of technology.
3. Innovation: Addresses a problem that has not yet been tackled (or has proven difficult to solve).
4. Impact: The number of people the technology potentially touches and the importance of its influence on their lives.
5. Representativity: Centrality to the topics covered by the conference (in this case ICME)
6. Magic: How closely the technology fills the description, "Any sufficiently advanced technology is indistinguishable from magic" (cf. http://en.wikipedia.org/wiki/Clarke%27s_three_laws)

Number 6 is basically a wild card that makes it possible to introduce in a controlled way that factor of je ne sais quoi, which seems to slip into the considerations made when judging demos in any case.

In practice, another factor that always seems to be important is how close the demo is to a working system that is or is about to be deployed in the real world. Also, when judging demos it seems that one is always trying to project forward: how important will this technology be five or ten years from now? Will the passage of time reveal that it is a disruptive technology? (Or, as Wikipedia prefers to call it disruptive innovation?)

For a list of the demos to be presented at ICME 2012, see the ICME 2012 Demo page.

I am dating the post 1 June, when I formulated the list of criteria. It's later now, but time seems to have simply gotten away from me, not surprising given the ICME 2012 Time Machine Session.