Showing posts with label MediaEval. Show all posts
Showing posts with label MediaEval. Show all posts

Sunday, February 26, 2017

Shared-tasks for multimedia research: Bans, benchmarks, and being effective in 2017

Last week, I officially resigned from contributing as an organizer to TRECVid Video Retrieval Evaluation, which is sponsored by NIST, a US government agency in Gaithersburg, Maryland. In 2016, I was part of the Video Hyperlinking task, and contributed by defining the year's relevance criteria, creating queries, and helping to design the crowdsourcing-based assessment. It has been a very difficult decision, so I would like to record here in this blogpost why I have made it. 

Ultimately, we make such decisions ourselves, and everyone navigates these difficult processes alone. However, it takes a lot of time and energy to search for the relevant information, and to weigh the considerations. For this reason, I think that for some it may be helpful to know more details about my own process.

Benchmarking Multimedia Technologies
Since 2008, I have been involved in benchmarking new multimedia technology. Benchmarking is the process of systematically comparing technologies by assessing their performance on standardized tasks. The process makes it possible to quantify the degree to which one algorithm outperforms another. Quantification is necessary in order to understand if a new algorithm has succeeded in improving over the state of the art, defined by the performance of existing algorithms.

The strength of benchmarking lies in the degree to which a benchmark succeeds in achieving open participation. If a new algorithm is compared to some, but not all, existing algorithms, the results of the benchmark reflect less clearly a true improvement over the state of the art.

My emphasis in benchmarking is on tasks that focus on the human and social aspects of multimedia access and retrieval. In other words, I am interested in people producing and consuming video, images, and audio content in their daily lives, and how technology can create algorithms to give them back usefulness and value from these activities. It is difficult to pack these aspects into quantitative metrics, so I am also committed to research that develops new evaluation methodologies and new metrics, as well.

Due to this emphasis, it is not surprising that most of my contribution has been channeled through the MediaEval Benchmark for Multimedia Evaluation. (I coordinate the MediaEval "umbrella", which synchronizes the otherwise autonomous tasks.) However, the strength of the benchmarking paradigm is weakened if a single benchmark, with a limited spectrum of topics, becoming all-dominant. Instead, we need to act to prevent a single effort from "taking over the market". We need to work towards ensuring that a broad range of different types of problems are investigated by the research community. Fostering breadth means offering not only multiple tasks, but multiple benchmarks. This year, I am again involved in MediaEval, but also, as last year, in contributing to the organization of  the NewsREEL task at CLEF (where my role is to contribute to design, documentation, and reporting).

Open Participation in Benchmarks
Both MediaEval and CLEF are open participation benchmarks in three aspects:
  • First, anyone can propose a task (there is an open call for tasks). CLEF chooses its tasks by multi-institutional committee, cf. 2017 CLEF Call for Task Proposals. MediaEval also chooses its task by multi-institutional committee. However, the committee checks only for viability. The ultimate choice lies in the hands of all community members, including organizers and participants, cf. MediaEval 2017 Call for Task Proposals. The goal of an open call for tasks is to promote innovation---constantly evolving tasks prevent the community from "locking in" on certain topics, and becoming satisfied with incremental progress.
  • Second, anyone can sign up to participate. Participants submit working notes papers, which go through a review process (emphasizing completeness, clarity, and technical soundness). MediaEval and CLEF both publish open access working notes proceedings.
  • Third, for both MediaEval and CLEF, workshop registration is open to anyone, and requires only the payment of a fee to cover costs. For MediaEval, the fee covers the costs of the workshop, and also of hosting the website and organizer teleconferences. People/organizations contribute time to cover the rest of workshop organization.
Like MediaEval and CLEF,  TRECVid also pursues the mission of offering an open research venue. Historically, both TRECVid and CLEF grew from TREC (also, of course, organized by NIST) so the commitment to the common cause is unsurprising in this sense. However, TRECVid does not offer open participation in all three of the above aspects. Specifically, there is no publicly circulated call for task proposals, and the workshop is closed. (The stated policy is that the workshop is only open to task participants, and "to selected government personnel from sponsoring agencies and data donors", cf. TRECVid 2017 Call for Participation) Technically, TRECVid is not able to welcome all participants. The US does not maintain diplomatic relationships with Iran. US Government employees cannot answer email from Iran. It is important to understand that this is a historical challenge, and is not new with the current US Republican Administration.

Defining Priorities and Making Decisions
Considerations related to open participation made me hesitant to get deeply involved in TRECVid. However, over the years, I have been very open for exchange. TRECVid originally reached out to me to give an invited talk back in 2009, when MediaEval was still VideoCLEF. (There are some musings on my blog from that trip.)  The idea was to learn from each other. We hope this year to reciprocate with a TRECVid speaker at CLEF/MediaEval.

In 2016, I contributed to the Video Hyperlinking organization, since the move of Video Hyperlinking from MediaEval to TRECVid represented a spread of the emphasis on the human aspects of multimedia retrieval, and it was important to me to support that explicitly.

All and all, it has taken a lot of time to decide where to invest my resources in 2017 in order to most effectively support multimedia benchmarking efforts that provide venues that are open and therefore effective as benchmarks.

With the new Republican Administration in the US, two considerations grew to dominate my decision making process. The first is how to contribute to the movement whose goal is to demonstrate the relevance and importance of science to the public and to policy makers https://www.forceforscience.org TRECVid, by virtue of being a benchmark, is certainly on the forefront of this movement (just by doing the same thing it has done for years). We need to support our US-based colleagues in the efforts to be a force for science, and hope that they support us as well, if we land in a similar situation.

The second is how to react to the travel ban, which would prevent scientists of certain countries from entering the US. The first-order effects of the travel have been constrained by court rulings. However, the future plans of the administration are uncertain, and there is a range of second-order effects that a court cannot un-do, e.g., people self-selecting out of participation since they are worried about their visa's being held up by additional processing steps (and granted, for example, only after the workshop has occurred). These secondary effects effectively prevent people from attending a US-based event even though technically they may be able to get a visa.

We are not alone in our thinking, but we are guided by a large number of organizations who have issued a public statement on the importance of openness for science (Statement of the International Council for Science, Statement of American Association for the Advancement of Science) including professional organizations that we belong to (Statement of the ACM, Statement of ACM SIGMM, Statement of IEEE) and European universities (Statement of the European University Association, Combined statement of all the universities in the Netherlands, Statement of Radboud University).

There is much power in making an open statement of values---more than one might think. However, we should avoid assuming that statements are enough and that the situation will go back to where it was before the current Republican Administration. In other words, the days are gone in which we had to dedicate relatively less time in protecting and upholding the values of openness in science. Instead, we need to think explicitly about where our effort can be best dedicated in 2017.

TREC/TRECVid celebrated their 25th anniversary in 2016. The event has been a constant through many changes of US administration, and it is heartening that the 2017 event will look, from the inside at least, with all probability pretty much like all other events over the past 25 years.

However, 2017 is the first year where people will be in the streets, in the US and around the world, marching for science:  https://www.marchforscience.com. The large-scale sense of urgency tells us that 2017 is not just business as usual. For this reason, it is important in 2017 to reexamine the idea that the US should be such a strong attractor within the map of scientific research in the world.

On top of the merit and can-do attitude that attracts people from around the world to US institutions, we as scientists (because we study systems and networks) know that another force is at play. Specifically, we know that US institutions enjoy preferential attachment, meaning that past success is a determiner of future success. This effect translates into the reality that new or small events (e.g., research topics or benchmarking workshops) need a lot of extra time and attention to establish or maintain themselves in the field. 2017 is the year that we need to think carefully about to which extent we want to contribute this non-linear feedback loop that strengthens the pull towards US-based events, and to which extent we want to build counterweights.

I consciously use the word "counterweights" since I am referring to a balancing act. We stand in complete solidarity with our US-based colleagues. Providing counterweights in no way detracts from that fact. For multimedia research, counterweights include region-based initiatives, and benchmarks that allow anyone to propose a task. A network of diverse benchmarks makes benchmarking as a whole stronger, and makes us internationally more robust,

My personal decision is that time spent promoting and preserving diversity is, in 2017, a more effective way to achieve the larger goals of benchmarking, than time spent reinforcing the connection between benchmarking and Gaithersburg, Maryland. I was born in Maryland, outside of DC, but Maryland is not where I am needed now. TRECVid will be fine without extra help from Europe, but what can (and does) suffer is the availability to the research community of non-US-based benchmarks.

Recommendations to TRECVid
The intention is for my resignation to be a positive decision for and not a negative decision against. Reasoning that my reflections on the topics are probably helpful to NIST, I distilled my thinking into a set of three recommendations. Interestingly, these recommendations are relatively independent of the situation in the US caused by the current Republican Administration:
  • First, TRECVid is an open research venue. I recommend stating this explicitly on the website. An example is the ACM Open Participation statement. 
  • Second, TRECVid is supported by NIST. I recommended a clearer statement of the source and the distribution of the funding on the website. People familiar with the benchmark know that NIST is the powerhouse behind its success, but it is not clear to newcomers. Critically, currently, the cases in which defense funding supports TRECVid are not clear. This is important to people who personally, or whose institutions, have a commitment to pursue research for civil purposes only. For example, many German institutions have a Zivilklausel by which they commit themselves to pursuing exclusively research for civilian purposes. Even if participation is nominally open, unclarity on defense funding can scare people away, and the benchmark is effectively not as open as it would otherwise aspire to be. (For completeness: at least one colleague assumed I received NIST funding for my work on Video Hyperlinking. I did not. The unclarity in the funding causes confusion.) 
  • Third, attention should be devoted to the archival status of the proceedings. As a good next step, they should be indexed by mainstream search engines. Moving forward, attention should be paid to maintaining a historical record of TRECVid should at some point in the future NIST not be able to continue to support open participation/open access in the way it does now.
If you have read all the way to the end of this blog post, let me finish by thanking you: both  for your dedication to open participation in scientific research, which is so essential to benchmarking, but also for taking the time to read about my personal struggle. It has been a long path.

Don't miss the March for Science on 22 April. Inspire and be inspired.

Or find another march around the world here: https://www.marchforscience.com/satellite-marches

Wednesday, February 8, 2017

Bytes and pixels meet the challenges of human media interpretation

Back in June, I gave a talk at the  Communication Science Department here at Radboud University Nijmegen. Today, I presented a version of that talk to my colleagues in the Language and Speech Technology Research Meeting. The abstract is below together with the slides, which are on SlideShare. During the discussion it became clear that many problems in natural language processing and information retrieval face the issue of human interpretations. It is important to find ways to move forward, although it may not be possible to pack our challenges into neat classification or ranking problems with a single set of consensus ground truth labels. A way forward, is to look to other disciplines for theory of how people understand and use media, and let these inform what we design our systems to do and the ways that we measure success.

Within computer science, "Multimedia" is a field of research that investigates how computers can support people in communication, information finding, and knowledge/opinion building. Multimedia content is defined broadly. It includes not only video, but also images accompanied by text and other information (for example, a geo-location). It can be professionally produced, or generated by users for online sharing. Computer scientists historically have a “love-hate” relationship with multimedia. They “love” it because of the richness of the data sources and the wealth of available data, which leads to interesting problems to tackle with machine learning. They “hate” it because multimedia is a diffuse and moving target: the interpretation of multimedia differs from person to person, and changes over time in the course of its use as a communication medium. This talk gives a view onto ongoing research in the area of multimedia information retrieval algorithms, which help people find multimedia. We look at a series of topics that reveal how pattern recognition, text processing, and crowdsourcing tools are used in multimedia research, and discuss both their limitations and their potential.


Sunday, April 3, 2016

Starting to RUN

Thank you for the email, tweets and texts about my new appointment at Radboud University Nijmegen. I'm happy that other people realize what a special day it was for me, and share my excitement about new opportunities and new challenges. I appreciate the warm reception at Radboud University. The "Welcome!" was unmistakeable: actually written on my whiteboard, when I walked into my office in the Center for Language Studies for the first time.

My appointment is as "Professor of Multimedia Information Technology" at the Faculty of Science, Institute for Computing and Information Sciences (iCIS). It involves a double affiliation (50/50) between iCIS and the Faculty of Arts, Centre for Language Studies (CLS). In this way, it brings together my background (pre-1990 in Math and EE; 1990-2000 in Formal Linguistics; and since 2000 in Computer Science, i.e., audio-visual search engines). It is a natural extension of this background that I will be working to bridge the research occurring on information access between the two faculties.

A press release about my appointment appeared on 31 March on the Radboud University homepage. I was very happy about the publicity for the MediaEval Multimedia Evaluation Benchmark. MediaEval is an initiative aimed at driving the development of new multimedia access technologies by offering shared tasks to the community. Instead of being centrally organized, it is grassroots in nature. My role is the bass player who, in a band, helps to links different parts together and keep the music moving forward on tempo. The success of the benchmark comes from the dedication and efforts of the task organizers, and the participants. (MediaEval is offering a great lineup of tasks in 2016, and signup is now open on the MediaEval 2016 website. The MediaEval 2016 workshop will be held 20-21 October 2016, right after ACM Multimedia 2016 in Amsterdam.)

Starting January 2017, Radboud University will be my main university (4 days per week), but I will maintain an affiliation with Delft University of Technology (1 day a week).

Currently, my main affiliation remains the Multimedia Computing Group at Delft University of Technology. However, I am at Radboud University Nijmegen for two days a week to get started at CLS. My first act is teach Intelligent Information Tools, a course for first and second year undergraduate students in Communication and Information Science. The students learn about the nature of information, the structure of the internet, how search, recommendation, and other information tools work, and also how to think critically about these tools.

At TU Delft I continue teaching, and pursuing my research. The main focus of my research at this time is recommender systems, within the context of the EC FP7 project CrowdRec "Fusion of active information for next generation recommender systems". It is a privilege to serve the CrowdRec consortium as the scientific coordinator.  Current highlights are: The NewsREEL news recommendation challenge, at CLEF 2016 the ACM RecSys 2016 job recommendation challenge, and the Workshop on Deep Learning for Recommender Systems, also at ACM RecSys 2016. I look forward to a successful conclusion of the project September 2016, and also to future collaborations.

Seven years ago, nearly to the day, I wrote the first post on this blog. I had read an article advising kill your blog, as an answer to blogposts getting lost in a sea of mainstream information. My post points out that it is strange to suggest that bloggers must change, and not mention the role or responsibility of search engines.

Now, I am more convinced in ever of the value of information within small circles. Search needs to support exploitation of that value. The readership of this blog is intended to be future versions of myself, and also a limited number of people interested in a deep dive into reflections on various search-related topics. As I move to a new university, and the number of people I teach or collaborate with grows, I would like to remember that. I'll probably have less time to write blog posts, but I have decided that I will wait a few more years until moving away from occasionally blogging.

Creating information is a way in which we help ourselves think. Intense conversations also refine thought. But the model of everyone talks to everyone about everything does not always make sense. Instead, we need room for reflection with a relatively small set of individuals. Search should support that.

What's blocking the road? Maybe we feel that small scale search is a success because Google now displays calendar events in our search results. Maybe facing the personal is somehow more laborious or painful. In any case, currently we are far from understanding the aggregated impact of thousands of local dialogues, or to evaluating the success of small search that helps us exchange ideas with our past selves, and our closest colleagues. The future holds no lack of challenges.


Saturday, March 5, 2016

A Non Neural Network algorithm with "superhuman" ability to determine the location of almost any image

Martha Larson and Xinchao Li

We would like to complement the MIT Technology Review headline Google Unveils Neural Network with “Superhuman” Ability to Determine the Location of Almost Any Image with information about NNN (Non Neural Network) approaches with similar properties.

This blogpost provides a comparison between the DVEM (Distinctive Visual Element Matching) approach, introduced by our recent arXiv manuscript (currently under review): 

Xinchao Li, Martha A. Larson, Alan Hanjalic Geo-distinctive Visual Element Matching for Location Estimation of Images (Submitted on 28 Jan 2016) (http://arxiv.org/abs/1601.07884)

and the PlaNet approach, introduced by the arXiv manuscript covered in the MIT Technology Review article:

Tobias Weyand, Ilya Kostrikov, James Philbin PlaNet—Photo Geolocation with Convolutional Neural Networks (Submitted on 17 Feb 2016) (http://arxiv.org/abs/1602.05314)

We also include, at the end, a bit of history on the problem of automatically "determining the location of images",  which is also known as geo-location prediction, geo-location estimation as in [3], or, colloquially, "placing" after [4].

Our DVEM approach is a search-based approach to the prediction of the geo-location of an image. Search-based approaches consider the target image (the image whose geo-coordinates are to be predicted) as a query. They then carry out content-based image search (i.e., query-by-image) on a large training set of images labeled with geo-coordinates (referred to as the "background collection"). Finally, they process the search results in order to make a prediction of the geo-coordinates of the target image. The most basic algorithm, Visual Nearest Neighbor (VisNN), simply adopts the geo-coordinates of the image at the top of the search results list as the geo-coordinates of the target image.  Our DVEM algorithm uses local image features for retrieval, and then creates geo-clusters in the list of image search results. It adopts the top ranked cluster, using a method that we previously introduced [5, 6]. The special magic of our DVEM approach is the way that it reranks the clusters in the results list: it validates the visual match at the cluster level (rather than at the level of an individual image) using a geometric verification technique for object/scene matching we previously proposed in [7], and it leverages the occurrence of visual elements that are discriminative for specific locations.

The PlaNet approach divides the surface of the globe into cells with an algorithm that adapts to the number of images in its training set that are labeled with geo-coordinates for that location, i.e., a location that has more photos will be divided into finer cells. Each cell is considered a class, and is used to train a CNN classifier.

Further comparison of the way the algorithms were trained and tested in the two papers:


DVEMPlaNet
Training set size5M images train, 2K validation91M train, 34M validation
Training set selectionCC Flickr images with geo-locations, (MediaEval 2015 Placing Task)Web images with Exif geolocations
Training time1 hour on 1,500 cores for 5M photos for indexing and feature extraction2.5 months on 200 CPU cores
Test set sizeca. 1M images2.3M images
Test set selectionCC Flickr images (MediaEval 2015)Flickr images with 1-5 tags
Train/test de-duplicationtrain/test sets mutually exclusive wrt uploading userCNN trained on near-duplicate images
Data set availabilityvia MM Commons on AWSnot specified
Model size100GB for 5M images377MB
BaselinesGVR [6], MediaEval 2015 IM2GPS [8]

From this table, we see that the training and test data for the algorithms are different, and for this reason, we cannot compare the accuracy measured for the two approaches directly. However, the numbers at the 1 km level (i.e., street level) suggest that DVEM and PlaNet are playing in the same ballpark. PlaNet reports correct prediction for 3.6% of the images on the (2.3M image test set) and 8.4% on the IM2GPS data set (237 images). Our DVEM approach achieves around 8% correct predictions on our 1M image test set, and is surprisingly robust to the exact choice of parameters. DVEM gains 12% relative performance over VisNN, and 5% over our own previous GVR. Note that [6] provides evidence that GVR outperforms IM2GPS [8]. PlaNet also reports that it outperforms IM2GPS, but the numbers are not directly comparable because 14x less training data is used.

The downside of search-based approaches is prediction time, as pointed out by the PlaNet authors in discussion IM2GPS. DVEM requires 88 hours on a Hadoop based cluster containing 1,500 cores to make predictions for 1M images. For applications requiring offline prediction, this may be fine, however, we assume that online geo-prediction is also important. We point out that with enough memory or an efficient index compression method, we would not need Hadoop, and we would be able to do the prediction on a single core with about 2s per query. Further, the question of how runtime scales is closely related to the question of the number of images that are actually needed in the background collection. Our DVEM approach uses 18x less training data than the PlaNet algorithm: if we are indeed in the same ballpark, this result calls in to question the assumption that prediction accuracy will not saturate after a certain number of training images.

We mention a couple reasons for which DVEM might ultimately turn out to out-perform PlaNet. First, the PlaNet authors point out that the discretization hurts accuracy in some cases. DVEM, in contrast, creates candidate locations "on the fly". As such, DVEM has the ability to make a geo-prediction at an arbitrarily small geo-resolution.

Second, the test set used to test DVEM is possibly more challenging than the PlaNet test set because it does not eliminate images without tags. We assume that the presence of a tag is at least a weak indicator of care on the part of the user. A careless user might also engage in careless photography, producing images that are low quality and/or are not framed to clearly depict their subject matter. A test set containing images taken by relatively more careful users could be expected to yield a higher accuracy.

Third, we assume that when near duplicates were eliminated from the PlaNet test/training set, that these were near duplicates from the same location. Eliminating images that are very close visual matches with other locations would, of course, artificially simplify the problem. However, it may also turn out that the elimination artificially makes the problem more difficult. In real life, a lot of people simply do take the same picture, for example, of the leaning tower of Pisa. A priori it is not clear how near duplicates should be eliminated to ensure the testing setup maximally resembles an operational setting.

The PlaNet paper was a pleasure to read, the name "PlaNet" is truly cool, and we are enthused about the small size of the resulting model. We are interested by the fact that "PlaNet" produces a distributional probability over the whole world, although we also remark that, DVEM is capable of producing top-N location predictions. We also liked the idea of exploiting sequence information, but think that considering temporal neighborhoods rather than temporal sequences might also be helpful. Extending DVEM with either temporal sequences or neighborhoods would be straightforward.

We hope that the PlaNet authors will run their approach using the MediaEval 2015 Placing Task data set so that we are able to directly compare the results. In any case, they will want to revisit their assertion that "...previous approaches only recognize landmarks or perform approximate matching using global image descriptors" in the light of the MediaEval 2015 Placing Task results, including our DVEM algorithm.

We would like to point out that work on algorithms able to predict the location of almost any image has been ongoing in full public visibility for a number of years. (Although given our field, we also enjoy the delicious jolt of a headline beginning "Google unveils...") The starting point can be seen as Mapping the World's Photos [9] in 2009. The MediaEval Multimedia Evaluation benchmark has been developing solutions to the problem since 2010, as chronicled in [10]. The most recent contribution was the MediaEval 2015 Placing task [11], cf. the contributions that use visual approaches to the task [12,13]. The MediaEval 2015 data set is part of the larger, publicly available YFCC100M data set, part of Multimedia Commons, and recently featured in Communications of the ACM [14]. MediaEval 2016 will offer a further edition of the Placing Task, which is open to participation for any research team who signs up.

We close by retuning to comment on the importance of  NNN (Non Neural Network) approaches. This example of the strength of DVEM vs. PlaNet provides a demonstration that there is reason for the research community to retain a balance in their engagement in NN and NNN approaches. One appealing aspect  of NNN approaches, and, in particular of search-based geo-location prediction, is the relative transparency of how the data is connected to the prediction. It may sound like science fiction from today's perspective, but one could imagine a future in which the person who took the image would receive a micro fee every time their image was used for the purpose of predicting geo-location metadata for someone else. Such a system would encourage people to take images that were useful for geo-location, and move us forward as a whole.

We would like to thank the organizers of the MediaEval Placing task for making the data set available for our research. Also a big thanks to SURF SARA for the HPC infrastructure without which our work would not be possible.

[1] Xinchao Li, Martha A. Larson, Alan Hanjalic Geo-distinctive Visual Element Matching for Location Estimation of Images (Submitted on 28 Jan 2016) (http://arxiv.org/abs/1601.07884)

[2] Tobias Weyand, Ilya Kostrikov, James Philbin PlaNet—Photo Geolocation with Convolutional Neural Networks (Submitted on 17 Feb 2016) (http://arxiv.org/abs/1602.05314)
[3] Jaeyoung Choi and Gerald Friedland. 2015. Multimodal Location Estimation of Videos and Images. Springer Publishing Company, Springer.
[4] P. Serdyukov, V. Murdock, R. van Zwol. 2009. Placing Flickr photos on a map. In Proceedings of the 32nd International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '09), ACM, New York, pp. 484–491.
[5Xinchao Li, Martha Larson, and Alan Hanjalic. 2013. Geo-visual ranking for location prediction of social images. In Proceedings of the 3rd ACM conference on International conference on multimedia retrieval (ICMR '13). ACM, New York, NY, USA, 81-88. 
[6Xinchao Li, Martha Larson, and Alan Hanjalic. Global-Scale Location Prediction for Social Images Using Geo-Visual Ranking, in IEEE Transactions on Multimedia, vol. 17, no. 5, pp. 674-686, May 2015.
[7] Xinchao Li, Martha Larson, Alan Hanjalic. 2015. Pairwise Geometric Matching for Large-scale Object Retrieval. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR '15), pp. 5153-5161.
[8] J. Hays and A. A. Efros, "IM2GPS: estimating geographic information from a single image," Computer Vision and Pattern Recognition, 2008. CVPR 2008. IEEE Conference on, Anchorage, AK, 2008, pp. 1-8.
[9] David J. Crandall, Lars Backstrom, Daniel Huttenlocher, and Jon Kleinberg. 2009. Mapping the world's photos. In Proceedings of the 18th international conference on World wide web (WWW '09,) ACM, New York, 761-770.
[19Martha Larson, Pascal Kelm, Adam Rae, Claudia Hauff, Bart Thomee, Michele Trevisiol, Jaeyoung Choi, Olivier Van Laere, Steven Schockaert, Gareth J.F. Jones, Pavel Serdyukov, Vanessa Murdock, Gerald Friedlan. 2015. The Benchmark as a Research Catalyst: Charting the Progress of Geo-prediction for Social Multimedia. In [3]. 
[11Jaeyoung Choi, Claudia Hauff, Olivier Van Laere, Bart Thomee. The Placing Task at MediaEval 2015. In Proceedings of the MediaEval 2015 Workshop, Wurzen, Germany, September 14-15, 2015, CEUR-WS.org, online ceur-ws.org/Vol-1436/Paper6.pdf
[12] Lin Tzy Li, Javier A.V. Muñoz, Jurandy Almeida, Rodrigo T. Calumby, Otávio A. B. Penatti, Ícaro C. Dourado, Keiller Nogueira, Pedro R. Mendes Júnior, Luís A. M. Pereira, Daniel C. G. Pedronette, Jefersson A. dos Santos, Marcos A. Gonçalves, Ricardo da S. Torres. RECOD @ Placing Task of MediaEval 2015. In Proceedings of the MediaEval 2015 Workshop, Wurzen, Germany, September 14-15, 2015, CEUR-WS.org, online ceur-ws.org/Vol-1436/Paper49.pdf
[13] Giorgos Kordopatis-Zilos, Adrian Popescu, Symeon Papadopoulos, Yiannis Kompatsiaris. CERTH/CEA LIST at MediaEval Placing Task 2015. In Proceedings of the MediaEval 2015 Workshop, Wurzen, Germany, September 14-15, 2015, CEUR-WS.org, online ceur-ws.org/Vol-1436/Paper58.pdf
[14] Bart Thomee, David A. Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, Li-Jia Li. YFCC100M: The New Data in Multimedia Research. Communications of the ACM, Vol. 59 No. 2, Pages 64-73.

Sunday, February 7, 2016

MediaEval 2015: Insights from last year's experiences in multimedia benchmarking



This blogpost is a list of bullet points concerning MediaEval 2015. It represents the "meta-themes" of MediaEval that I perceived to be the strongest during the MediaEval 2015 season, which culminated with the MediaEval 2015 Workshop in Wurzen, German (14-15 September 2015). I'm putting them here, so we can look back later and see how they are developing.
  1. How not to re-invent the wheel? Providing task participants with reading lists of related work and with baseline implementations helps ensure that it is as easy as possible for them to develop algorithms that extend the state of the art.
  2. Reproducibility and replication: How can we encourage participants to share information about their approaches so that their results can be reproduced or replicated? How can we emphasize the importance of reproduction and replication and at the same time push for innovation, and forward movement in the state of the art (and avoid re-inventing the wheel as just mentioned)? One answer that arose this year was to reinforce student participation. Students should feel welcome at the workshop, even if they “just” reproduced an existing workflow.
  3.  Development of evaluation metrics for new tasks: Innovating a new task may involve a developing a new evaluation metric. All tasks face the challenges of ensuring that they are using an evaluation metric that faithfully reflects usefulness to users within an evaluation scenario.
  4. How to make optimal use of leaderboards in evaluation: Participants should be able to check on their progress over the course of the benchmark, and aspire to ever-greater heights. However, it is important that leaderboards not discourage participants from submitting final runs to the benchmark. It is possible that an innovative new approach does very badly on the leaderboard, but is still valuable.
  5. Understanding the relationship between the conceptual formulation of the task, and the dataset that is chosen for use in the task: Are the two compatible? Are there assumptions that we are making about the dataset that do not hold? How can we keep task participants on track: solving the conceptual formulation from the task, and not leveraging some incidental aspect of the dataset?
  6. Disruption: Tasks are encouraged to innovate from year to year. However, 2015 was the first year that organizers started planning far ahead for “disruption” that would take the task to the next level in the next year.
  7. Using crowdsourcing for evaluation: How to make sure that everyone is aware of and applies best practices? How to ensure that the crowd is reflective of the type of users in the use scenario of the task?
  8. Engineering: Task organization involves an enormous amount of time and dedication to engineering work. We continuously seek ways to structure organizer teams and to recruit new organizers and task auxiliaries to make sure that no one feels that their scientific output suffered in a year where they spend time handling the engineering aspects of MediaEval task organization.
  9. Defining tasks and writing task descriptions: We repeatedly see that the process of defining and new task and of writing task descriptions must involve a large number of people. If people with a lot of multimedia benchmarking experience contribute, they can help to make sure that the task definition is well grounded in the existing literature. If people with very little experience in multimedia benchmarking contribute, they can help to make sure that the task definition is understandable even to new participants. We try to write task descriptions such that a master student planning to write a thesis in a multimedia related topic would easily understand what was required for the task.

In order to round this off to a nice "10" points let me mention another issue that is constantly on my mind, namely, the way that the multimedia community treats the word "subjective".

"Subjective" is something that one feels oneself as a subject (and cannot be directly felt by another person---pain is the classic example). In MediaEval tasks, such as Violent Scene Detection, we would like to respect the fact that people are entitled to their own opinions about what constitutes a concept. Note that people can communicate very well concerning acts of violence, without all having an exactly identical idea of what constitutes "violence". Because the concept "works" in the face of the existence of person perspectives, we can consider the task "subjective". 

So often researchers reason in the sequence, "This task is subjective, therefore it is difficult for automatic multimedia analysis algorithms to address". That reasoning simply does not follow. Consider this example: Classifying a noise source as painful is the ultimate "subjective task". You as a subject are the only one who knows that you are in pain. However: Create a device that signals "pain" when noise levels reach 100 decibels, and you have a solution to the task. Easy as pie. "Subjective" tasks are not inherently difficult. 

Instead: whether a task is difficult to address with automatic methods depends on the stability of content-based features across different target labels. 

The whole point of machine learning is to generalize across not only obvious cases, but also across cases in which no stability of features is apparent to a human observer. If we stuck to tasks that "looked" easy to a researcher browsing through the data, (exaggerating a bit for effect) we might as well handcraft rule-based recognizers. So my point 10 is to try to figure out a way to keep researchers from being scared off from tasks just because they are "subjective", without giving the matter a second thought. Multimedia research needs to tackle "subjective" tasks in order to make sure that it remains relevant to the real-world needs of users---once you understand subjectivity, you start to realize that it is actually all over the place.

In 2014, we noticed that the discussion of such themes was becoming more systematic, and that members of the MediaEval community were interested in having a venue in which they could publish their thoughts. For this reason, in 2015, we added a MediaEval Letters section to the MediaEval Working Notes Proceedings dedicated to short considerations of themes related to the MediaEval workshop. The Letter format allows researchers to publish their thoughts already as they are developing, even before they are mature enough to appear in a mainstream venue.

The concept of MediaEval Letters was described in the following paper, in the 2015 MediaEval Working Notes Proceedings:

Larson, M., Jones, G.J.F., Ionescu, B., Soleymani, M., Gravier, G. Recording and Analyzing Benchmarking Results: The Aims of the MediaEval Working Notes Papers. Proceedings of the MediaEval 2015 Workshop, Wurzen, Germany, September 14-15, 2015, CEUR-WS.org, online http://ceur-ws.org/Vol-1436/Paper90.pdf


Look for MediaEval Letters to be continued in 2016.

Saturday, August 22, 2015

Choosing movies to watch on an airplane? Compensate for context

For some people, an airplane is the perfect place to catch up on their movies-to-watch list. For these people there is no difference between sitting on a plane and sitting on the couch in their living room.

If you are one of these people, you are lucky.

If not, then you make want to take a few moments to think about what kind of a movie you should be watching on an airplane.

These are our two main insights on how you should make this choice:
  • Watch a movie that uses a lot of closeups (or relatively little visual detail), is well lit, and moves relatively slowly so that you can enjoy it on a small screen.
  • Watch something that is going to engage you. Remember that the environment and the disruptions on an airplane might affect your ability to focus, and, in this way disrupt your ability to experience an empathetic relationship with the characters. In other words, unless the plot and characters really draw you in movie might not "work" in the way it is intended.
For an accessible introduction to how movies manipulate your brain see the Wired series on Cinema Science article:
http://www.wired.com/2014/08/how-movies-manipulate-your-brain
At the perceptual level your brain needs to be able exercise its ability to "stitch things together to make sense". It's plausible that this "stitching" also has to be able to take place at an emotional level. Certain kinds of distractors can be expected to simply get in the way of that happening as effectively as it is meant to.

When we began to study what kinds of movies that people watch on planes we used these two insights as a point of departure. We started with these insights after having made some informal observations about the nature of distractors on an airplane, which are illustrated by this video.



At the end of the video, we formulate the following initial list of distractors, which impact what you might want to watch on an airplane.
  1. Engine noise
  2. Announcements
  3. Turbulence
  4. Small screen 
  5. Glare on screen
  6. Inflight service
  7. Fellow travelers
  8. Kid next to you
In short, when choosing a movies to view on the airplane, you should pick a movie that can "compensate for context", meaning that you can enjoy it despite the distractors inherent in the situation aboard an airplane.

We are looking to expand this list of distractors as our research moves forward.

Our ultimate goal (still a long way off) is to build a recommender system that can automatically "watch" movies for you ahead of time. The system would be able to suggest movies that you would like, but above and beyond that the suggested movies would be prescreened to be suitable for watching on an airplane. Such a system would help you to quickly decide what to turn on at 30,000 feet, without worrying that half way through you will realize that it might not have been a good choice.

Since we started this research, I have been paying more an more attention to the experiences that I have with movies on a plane. Here are two.

  • On a recent domestic flight: The woman next to me started to watch Penny Dreadful. She turned it off about ten minutes in. I then also tried to watch it. I really like the show, but it's meant to be dark, gory, and mysterious. These three qualities turned into poorly visible, disconcerting, and confusing at 10,000 feet. This is what I am trying to capture in the video above.
  • On a recent Transatlantic flight: The man sitting next to me turned on his monitor, and turned on The Color Purple as if it were on his watch list. It's probably on most people's movies-to-watch list, so this seems like a safe choice. However, I was trying to review a paper, and was subject to over two hours of unavoidable glimpses of violence on a screen a few feet from my own. It's the emotional impact of these scenes that make it a great movie. By the same token, you might not want to be watching it on a plane, especially if you are not going to experience the entire emotional arc. (The movie should be watched with full focus, from beginning to end.)

Until now, the work on context-aware movie recommender systems that I have encountered has recommended movies for situations that are part of what is considered to be "normal" daily life, e.g.,  watching movies during the week vs. on the weekend, watching movies with your kids vs. with your spouse. We need more recommender system work that will allow us to get to movies that are suitable for less ideal, more unpleasant, perhaps less frequent situations.  Why waste a good movie by watching it in the wrong context? And why suffer anymore than necessary while on an airplane?

The work is being carried out within the context of the MediaEval Benchmarking Initiative for Multimedia Evaluation, see Context of Experience 2015. It owes a lot to the CrowdRec project, which pushes us to understand how we can make recommenders better by asking people explicitly to contribute their input.

Thursday, October 23, 2014

MediaEval 2014 Placing Task Technical Retreat

At the end of MediaEval Workshop, the 2014 Placing Task had a technical retreat where the details of the year's crop of algorithms was assessed, and plans for the future were discussed.

I missed the first part of the meeting, because I was still doing some organizational stuff, and also saying goodbye to people (I find that so difficult to do, and I certainly didn't manage to say goodbye to everyone). However, I did take some notes on the parts that I attended and I am putting them here for posterity. 

Of course expect the usual attentional bias in what I chose to write down---possibly also in the categories that I put the notes into as well.

Moving beyond geo-location estimation
Can we formulate a data analysis task that moves beyond geo-prediction?

Can we drive the benchmark to get the task participants to uncover the weaknesses in current placing systems. What mistakes are you making, are why are you making them?

Geo-relevance
  • For which images is geo-location relevant?
  • Which is the location for which it is relevant?
  • What is the tolerance for error? (depends on humans, applications)
Placability
In the past two years placability has been offered as part of the task, but has been disappointingly unpopular. This seems to be a matter of people not having time. We shouldn’t take the evidence as meaning that people don’t want to do it.

Alternate form of placability:
“Select a set of x images (e.g., 100 images) from the test set that you are sure that you have placed correctly and visualize them in a map"

How to support the participants
Can we release a baseline system?
Estimates for the error?

How to move beyond co-ordinate estimation
  • Can we make the Placing Task more clearly application oriented?
  • Are there use scenarios beyond Flickr?
  • Is anyone interested in the task of Geo-Cloaking? 
  • Can the task pit two teams  against each other, one cloaking and one placing?
Evaluation metric

  • We think that geodesic distance is convenient, but has limits, since it doesn’t reflect the usefulness of predictions for humans within use scenarios.
  • Maybe move to administrative districts
  • Other metrics motivated by human image interpretation?

Ground truth
We can measure placing performance only within the error of the ground truth (cf. [2]). What can we do to work around this limitation?
  • Correspondence between geo-tags and exif metadata is indicative of whether the tag is correct. See also cool new work on timestamps [4].
  • Are their other easily measurable characteristics of images online that can be used to identify images/videos with reliable geo-tags at a large scale?
  • Collect more human labeled data. Do we really need to have a 500,000 item size data set?
How People Judge Place
Users (i.e., humans judging images) have different ways of knowing where a picture was taken. 

It depends on the relationship between the human judging, the image, the act of image creation.

The most basic contrast is between the case in which the human judge is the photographer, and the case in which the human judge is not the photographer and also shares no life experiences with the photographer.

Previously I discussed these different relationships in a post entitled “Visual Relatedness is in the Eye of the Beholder” and also in [3].

Why is this important? Some mistakes that are made by automatic geo-location prediction algorithms are disturbing to users, some are not. Whether or not a mistake is disturbing to a particular human judge is related to the way in which the human judge knows where the picture was taken. In other words, I may “forgive” an automatic geo-location estimation algorithm for interchanging the location of two rock faces of the same mountain, unless one of them happens to be the rock face that I myself managed to scale. How people judge place, is closely related to the types of evaluation metrics we need to choose to make the Placing Task as useful as possible.

In the Man vs. Machine paper [1] sets up a protocol that gathers human judgements in a way that controls the way in which people “know” or are allowed to come to know the location of images. More work should be explicitly aware of these factors.

Embrace the messiness
The overall conclusion: anything that we can do to move the task away from "number chasing” towards insight is helpful. This means finding concrete ways to embrace the fact that the task is inherently messy.

Thank you!
Thank you to the organizers of Placing 2014 for their efforts this year. We look forward to a great task again next year.

References
[1] Jaeyoung Choi, Howard Lei, Venkatesan Ekambaram, Pascal Kelm, Luke Gottlieb, Thomas Sikora, Kannan Ramchandran, and Gerald Friedland. 2013. Human vs machine: establishing a human baseline for multimodal location estimation. In Proceedings of the 21st ACM international conference on Multimedia (MM '13). ACM, New York, NY, USA, 867-876. 
[2] Claudia Hauff. 2013. A study on the accuracy of Flickr's geotag data. In Proceedings of the 36th international ACM SIGIR conference on Research and development in information retrieval (SIGIR '13). ACM, New York, NY, USA, 1037-1040.
[3] M. Larson, P. Kelm, A. Rae, C. Hauff, B. Thomee, M. Trevisiol, J. Choi, O. van Laere, S. Schockaert, G. J. F. Jones, P. Serdyukov, V. Murdock, and G. Friedland. The benchmark as a research catalyst: Charting the progress of geo-prediction for social multimedia. In J. Choi and G. Friedland, editors, Multimodal Location Estimation of Videos and Images. Springer, 2015.
[4] Thomee, B., Moreno, J.G.,, Shamma, D.A. Who’s Time Is it Anyway? Investigating the Accuracy of Camera Timestamps. ACM MM 2014, to appear. http://www.liacs.nl/~bthomee/assets/14time_p.pdf

Sunday, October 27, 2013

Power behind MediaEval: Reflecting on 2013

Power behind MediaEval 2013
The MediaEval 2013 Workshop was held in the Barri Gòtic, the Gothic Quarter, of Barcleona on 18-19 October, just before ACM Multimedia 2013 (held in Barcelona). The workshop was attended by 100 participants, who presented and discussed the results that they had a achieved on 11 tasks, 6 main tasks, and on 5 "Brave New Tasks" that ran for the first time this year.

The workshop produced a working notes proceedings containing 96 short working papers, and has been published by CEUR-WS.org.

The purpose of MediaEval is to offer tasks to the multimedia community that  support research progress on multimedia challenges that have a social or a human aspect. The tasks are autonomous and each run by a different team of task organizers. My role within MediaEval is to guide the process of running tasks, which involves providing feedback to task organizers and sending out the cues that keep the tasks ruining smoothly and on time. Today I did a quick count that revealed that during the 2013 season, I wrote 1529 personal emails to people that contained the keyword "MediaEval" in them.

What makes MediaEval work, however, cannot be expressed in numbers. Rather, it is the dedication and intensive effort of a large group of people, who propose and organize tasks and carry out the logistics that make the workshop come together. My motivation to continue MediaEval year after year stems largely from an underlying sense of awe at what these people do: both at the work that I am aware of and also at the many things that they do behind the scenes that make largely invisible. These people are the power behind MediaEval. Here I represent them with the picture above, which are the power plugs arranged by Xavi Anguera from Telefonica with the assistive effort of Bart Thomee from Yahoo! Research. The process involved a combination of precision car driving and applied electrical engineering.

In the airplane back from Barcelona yesterday, I finished processing the responses that we received from the participant surveys (collected during the workshop), input from the organizers meeting (held on Sunday after the workshop), and feedback that people gave me verbally during ACM Multimedia (last week). These points are summarized below.

Thus endeth MediaEval 2013, but at the same time beginneth the season of MediaEval 2014. Hope to have you aboard.

Community Feedback from the MediaEval 2013 Multimedia Benchmarking Season + Workshop

The most important feedback point this year was the new structure of the workshop, which was very well received. This year the workshop was faster paced and we introduced poster sessions. We were happy that people liked the short talks and that the poster sessions were considered to be useful and productive. There is a clear trend to preferring there to be more discussion time at the workshop, both in the presentation sessions and in the poster sessions. An idea for the future is to separate passive poster time (posters are hanging and people can look at them but the presenter need not be present) from active poster time (presenter is standing at the poster).

The number one most frequent request was for MediaEval to provide more detailed information. This request was made with respect to a range of areas: descriptions of the tasks should always strive to be maximally explicit; descriptions of the evaluation methods should be detailed and available in a timely manner; task overview talks at the workshop should contain examples and descriptions that allow a general audience (i.e., people who did not participate in the task) to understand the task easily.

Other suggestions were to increase consistency check and continue to promote industry involvement. Finally, requests for more time for preparation of presentations and to explicitly invite (and support) groups to make demos with the posters.

The organizers meeting on Sunday was the source of additional feedback. Task organization requires a huge amount of time and dedication from task organization teams and it is important that this is distributed as evenly as possible across the year and across people. In general, tasks would benefit from additional practical guidance on organization. This includes task management and evaluation methodologies. Since MediaEval is a decentralized system, the source of this guidance must be people with past experience with task organization and communication between tasks. Here, the bi-weekly telcos for organizers are an important tool.

In the coming year, the awards and sponsorship committee can expect an expanded role. The outreach to early-career researchers and to researchers located outside of Europe (in the form of travel grants) is seen by the organizers to be not merely a "nice-to-have", but rather a central part of MediaEval's mission. There is solid consensus about the usefulness of the MediaEval Distinctive Mentions (MDMs). MDMs are peer-to-peer certificates awarded by task organizers to each other or to the participants of their tasks. The MDMs  allow the community to send public messages between members of the community, and especially to point out participant submissions that are highly innovative or have particularly high potential (although they may not have been top scorers according to the official evaluation metric). It is important to make clear that the MediaEval Distinctive Mention is not an "award", since the process by which they are chosen is intentionally kept very informal. In the coming year, we will be investigating the issue of whether MediaEval should introduce a five-year impact award, that would be more formal in nature. The peer-to-peer MDMs will be maintained, although and effort will be made to make them increasingly transparent.

In general we were satisfied with the process used to produce the proceedings. Having groups do an online check of their metadata was helpful. If future years also involve proceedings with 50+ papers, we will need to further streamline the schedule for submission---with the ultimate goal of having the proceedings online at the moment that the workshop opens.

Saturday, July 13, 2013

Event Detection in Multimedia: Different definitions, different research challenges.

Yesterday, someone asked me for a pointer to work in the area of event detection in multimedia content. This mail prompted me to finally get out a blog post that explains the distinction between the different sorts of underlying challenges that researchers are referring to when they discuss events in multimedia.

A simple definition of an event is a triple (t, p, a), consisting of a moment a time t, a place in space p, and one or more actors a. For example,  at (t=) 2pm 13 July 2013. at the (p=) Faculty of Electrical Engineering, Mathematics and Computer Science at Delft University of Technology, (a=) I am now involved in a "blog-post writing" event.

Let's look at that definition of event in terms of aspects that matter to us as multimedia researchers. If you took a video of me (here and now), another human looking at your video might notice that I am also eating a salad. Consequently, the video could equally be considered to depict a "lunch-eating event". For this reason, it makes sense to also introduce a fourth variable "v", to arrive at (t, p, a, v). The "v"stands for the name for the action or the activity part of the event. I use "v" for "verb" since these names corresponds to verbs or can be expressed by phrases involving verbs.

Note that there are at least three basic ways to name (or label) events when it comes to multimedia: (1) name the event from the perspective of the/an actor (I, the actor, call it a lunch eating event, because I know it is lunch.) (2) name the event from the perspective of the person recording the multimedia (The person sees me engaged in an eating event, but do not necessarily know or care that it is lunch.) and (3) name the event from the perspective of the/a person looking at the multimedia in a time and place other than when and where it was captured (The person sees me sitting at a computer, but does not notice or want to pay attention to the salad.) Many times these three perspectives collapse and there is only a single label that would be relevant, but it should be kept in mind that they do not necessarily do so. We risk over-simplifying the world and losing valuable information if we assume that they can be conflated. Instead, multimedia systems must be careful to maintain multiple views, i.e., a video that for one person (e.g., a government official) depicts a riot, might for another (e.g., a concerned citizen) depict a demonstration.

The (t, p, a, v) definition of an event is sometimes constrained by a fourth factor, namely, advanced human planning. Multimedia research that looks at planned events focuses on events that humans organize for social purposes and that can therefore be anticipated in advance of their occurrence. This group of events includes events like concerts, games, conferences and parties. It is generally referred to as "Social Event Detection".

The Social Event Detection (SED) task at the MediaEval Multimedia Benchmark started in 2011 and has been drawing a steadily increasing number of participants each year.  MediaEval SED 2013 offers the most ambitious and interesting SED task to date. The SED Task Organizers have organized workshops and special sessions at various conferences, for example, recently the Special Session on Social Events in Web Multimedia at ICMR 2013. The MediaEval bibliography includes a relatively up-to-date list of the papers that have been published regarding the MediaEval SED task.

The SED task is defined such that its multimedia aspect arises because addressing the task requires combining different information sources (text, photos, videos) from different social communities on the Web. Note that it is the use of the (t, p, a, v) definition of an event and not per se the social nature of the data that distinguishes SED from other types of event detection in multimedia.

Another important type of event detection is defined as involving not the full (t, p, a, v), but rather (v). In other words, this variant of event detection is interested not in any specific event, but in detecting the occurrence of instances of a particular event type. This type of event detection is referred to as Multimedia Event Detection and has been offered as a task in TRECVid since 2013.  Examples of these sorts of events are "Birthday party" (from TRECVid MED 2011) and "Giving directions to a location" and "Winning a race without a vehicle" (from TRECVid MED 2012).

If you consider only the labels that they use to refer to events, SED and MED look very much the same. However, it is important to remember that for MED, multimedia that is considered relevant to the event "birthday party" can depict any birthday party at any time, at any place around the world. In other words, for MED "birthday party" is an instantiation of any event of the type birthday party. Only (v) and not the full (t, p, a, v) are part of the definition. For SED, "birthday party" would be for example, my birthday party, taken on my birthday in 2013, at the particular place at which I celebrated.

Again make note of the task definition of MED. The MED task is defined such that its multimedia aspect arises because addressing the task requires combining different modalities within the same video (visual + audio channel). Typically, the data is not social video per se. Note that it is the use of the (v) definition of an event and not per se the nature of the data that distinguishes MED from other types of event detection in multimedia.

My impression is that some researchers in the community are convinced that the way forward for the research community is to first detect (v) (i.e., apply the MED event definition) and then filter the detection results to be constrained to (t, p, a, v) (i.e., generate a list of results that follows the SED definition).

Before making this assumption, I would urge researchers to carefully contemplate the use scenario of their applications. For example, if I have one picture from a birthday party I attended and I want to search for other pictures of the same birthday party on the Internet, it does not make any sense at all to solve the complete MED birthday party detection problem as the first step in the process.

As far as weddings go, if we are using visual features it's tempting to rely on the presence of that beautiful white wedding gown to detect instances of v = "wedding". However, despite an initial impression that the gown provides a stable visual indicator, its not going to get you very far in a  real-world data set:

Leaving courthouse on first day of gay marriage in Washington
Ultimately, we need to look at both (t, p, a, v) and (v) and all the definitions of event detection in multimedia that lie in between. Luckily there are events in which researchers with all perspectives come together. I have now finished both my lunch and my blog post and, as a final note, I leave you with an example of just such an event:

Vasileios Mezaris, Ansgar Scherp, Ramesh Jain, Mohan Kankanhalli, Huiyu Zhou, Jianguo Zhang, Liang Wang, and Zhengyou Zhang. 2011. Modeling and representing events in multimedia. In Proceedings of the 19th ACM international conference on Multimedia (MM '11). ACM, New York, NY, USA, 613-614

Tuesday, April 30, 2013

The Five Runs Rule: Less is More in Multimedia Benchmarking

MediaEval is a multimedia benchmarking initiative that offers tasks in the area of multimedia access and retrieval that have a human or a social aspect to them. Teams sign up to participate, carry out the tasks, submit results, and then present their findings at the yearly workshop.

I get a lot of questions about something in MediaEval that is called the "five-runs rule". This blogpost is dedicated to explaining what it is, where it came from, and why we continue to respect it from year to year.

Participating teams in MediaEval develop solutions (i.e., algorithms and systems) that address MediaEval tasks. The results that they submit to a MediaEval task are the output generated by these solutions when they are applied to the task test data set. For example, this output might be a set of predicted genre labels for a set of test videos. A complete set of output generated by a particular solution is called a "run". You can think of a run as the results generated by an experimental condition. The five-run rule states that any given team can only submit of five sets of results to a MediaEval task in any given year.

The simple answer to why MediaEval tasks respect the five-runs rule is "They did it in CLEF'. CLEF is the Cross Language Evaluation Forum, now the Conference and Labs of the Evaluation Forum, cf. http://www.clef-initiative.eu/. MediaEval began as a track of CLEF in 2008 called VideoCLEF and at that time we adopted the five-runs rule and have been using it ever since.

The more interesting answer to why MediaEval tasks respect the five-runs rule is "Because it makes us better by forcing us to make choices". Basically, the five-runs rule forces participants during the process of developing their task solutions to think very carefully about what the best possible approach to the problem would be an focus their effort there. The rule encourages them to use a development set (if the task provides one) in order to inform the optimal design of their approach and select their parameters.

The five-runs rule discourages teams from "trying everything" and submitting a large number of runs, as if evaluation was a lottery. Not that we don't like playing the lottery every once in a while, however, if we choose our best solutions, rather than submitting them all, we help to avoid over-fitting new technologies that we develop to a particular data set and a particular evaluation metric. Also, if we think carefully about why we choose a certain approach when developing a solution, we will have better insight about why the solution worked or failed to work...which gives us a clearer picture of what we need to try next.

The practical advantage of the five-runs rule is that it allows the MediaEval Task Organizers to more easily distill a "main message" from all the runs that are submitted to a task in a given year: the participants have already provided a filter by submitting only the techniques that they find most interesting or predict will yield the best performance.

The five-runs rule also keeps Task Organizers from demanding too much of the participants. Many tasks discriminate between General Runs ("anything goes") and Required Runs that impose specific conditions (such as "exploit metadata" or "speech only"). The purpose of having these different types of runs is to make sure that there are a minimum number of runs submitted to the task that investigate specific opportunities and challenges presented by the data set and the task. For example, in the Placing Task, not a lot of information is to be gained from comparing metadata-only runs directly with pixel-only runs. Instead, a minimum number of teams have to look at both types of approaches in order for us to learn something about how the different modalities contribute to solving the overall task. Obviously, if there are two many required runs, participating teams will be constrained in the dimensions along which they can innovate, and that would hold us back.

Another practical advantage of the five-runs rule has arisen in the past years when tasks, led by Search and Hyperlinking and also Visual Privacy, have started carrying out post hoc analysis. Here, in order to deal with the volumes of the runs that need to be reviewed by human judges (even if we exploit crowdsourcing approaches), it is very helpful to have a small focused set of results.

Many tasks will release the ground truth for their test sets at the workshop, so teams that have generated many runs can still evaluate them. We encourage the practice of extending the two-page MediaEval working notes paper into a full paper for submission to another venue, in particular international conferences and journals. In order to do this, it is necessary to have the ground truth. Some tasks do not release the ground truth for their test sets because the test set of one year becomes the development set of the next year, and we try to keep the playing field as level as possible for new teams that are joining the benchmark (and may not have participated in previous year's editions of a task).

In the end, people generally have the experience that when they are writing up their results in their two-page working notes paper, and trying to get to the bottom of what worked well and why in your failure analysis, they are generally quite happy that they are dealing with not more than five runs.

Monday, March 12, 2012

The MediaEval Workshop: What it's meant to be and why you want to be there.

The MediaEval workshop is the event held each year in the fall at the culmination of the yearly benchmarking cycle. At the time of the MediaEval workshop many things have already happened in the benchmarking year: The task organizers have worked hard to define tasks and issue data sets. Participants have worked hard to develop algorithms that tackle the tasks, and they have run these algorithms on the data sets. The "runs" have been evaluated and each participating team has written their working notes paper. Now it's time for the workshop!

This blog post provides a view (from my perspective as one of the MediaEval co-ordinators) on the history of the workshop and on what the workshop is meant to be. In particular, it highlights the similarities and differences with other types of workshops.

The main goal of the MediaEval workshop is to bring everyone who carried MediaEval tasks together in one physical location to present and discuss their results, exchange experiences and develop ideas for how to improve their algorithms. The first year that we met at a medieval convent, Santa Croce in Fossabanda, it was mostly due to delight in the wordplay between medieval and MediaEval. However, we soon came to appreciate how getting everyone together working, eating and basically living in the same space creates an extremely productive focus on our our common tasks and goals. (Although unlike the nuns of the Middle Ages we don't go dashing off to prayer when we hear the convent bell.)

In 2010, we held the MediaEval workshop just before ACM Multimedia 2010, which was held at the Palazzo dei Congressi in Firenze, Italy. Santa Croce in Fossabanda is located in Pisa, about an hour's train ride from Firenze. We chose the dates and place to cut down on travel time and cost for people who wanted to attend both ACM Multimedia 2010 and the MediaEval 2010 workshop.

The next year, Interspeech 2011 came to the Palazzo dei Congressi in Firenze about the same time of year. MediaEval submitted a proposal and was granted the status of an "Official Satellite Event of Interspeech 2011". Now, instead of just taken advantage of the convenience for travel, we began emphasize hook-up to the topic of the conference: being associated with Interspeech reinforced the use of speech within MediaEval and helped us to better realize the goals expressed in the MediaEval slogan The "multi" in multimedia.

This year, the 12th European Conference on Computer Vision (ECCV 2012) will be held in Firenze. This conference provides us with the opportunity to reinforce the use of visual content within MediaEval, another of the multimedia "multi's". The MediaEval 2012 Workshop will be held right before this conference starts, again in Santa Croce in Fossabanda in Pisa. A very close contending idea for the MediaEval 2012 workshop was to hold it near 13th International Society for Music Information Retrieval Conference (ISMIR 2012) in Porto. However, the results of the MediaEval 2012 survey showed that the majority preferred to co-ordinate the date and place with ECCV 2012 and stay for a third year in Pisa.

In sum, the idea of being close to a large conference related to MediaEval topics has grown from being a convenience to being an aspect that strengthens and enriches both the benchmark and the workshop.

What exactly happens at the MediaEval workshop? The workshop consists of a series of sessions on the individual tasks. The task organizers present the task as a whole and each team presents its individual results, after which the floor is opened for discussion. More discussion and exchange occurs during meals and breaks in ad hoc groups. We try to build in a lot of space for discussion into the workshop schedule and especially try to create opportunities for students to discuss with more experienced researchers who help them to guide their efforts along the most effective path.

The proceedings are an important aspect of the workshop. The MediaEval workshop proceedings is a "Working Notes Proceedings" and consists of short (two page) papers written by the participating teams to report their results. These papers describe the algorithms that are used and present the results. They then also seek to understand the algorithm, and participants are requested to report:
  • which cases are easy/difficult and why
  • which approaches work best and why
  • which approaches do not work well
The working notes submissions are reviewed by the task organizers. The task organizers may either accept the submission as is, or may come back to the participant team with a request for revisions of the paper. The preferred mechanism is to have the papers revised rather than to reject them --- sometimes this revision cycle means that the working notes proceedings is not ready until just before the workshop. For this reason, the working notes proceedings is distributed at the workshop on a memory stick.

After the workshop, the working notes is made available online. The 2011 the "Working Notes Proceedings of the MediaEval 2011 Workshop" was published here: http://ceur-ws.org/Vol-807/ and in the future we would like to continue to use http://ceur-ws.org/. By publishing in this way, the copyright for the individual papers is with the papers' authors.

MediaEval working notes papers are intended to be first versions of work that is later extended by the authors and submitted to mainstream venues, such as conferences and journals. The fact that the MediaEval workshop proceedings consists of short working notes and that the copyright does stays with the authors keep the proceedings consistent with the goal of reporting an initial research result, which will then be refined and extended using input from the discussion at the MediaEval workshop.

Another important goal of the workshop is to discuss the tasks themselves. Did they help to move the state-of-the-art forward? Should we improve or replace them next year? Are their new questions that need to be answered that require new tasks. Any one attending the workshop is welcome to stand up in the final session and "pitch" an idea for a new task. Tasks which receive good community support in this session have a good chance of receiving the response levels they need on the yearly MediaEval survey to run as tasks in the next year.

Finally, the workshop also aims to connect ourselves and our research to the larger community. We welcome participants from industry who have tasks that they might want to pitch to the community. We also welcome representatives from other benchmarks: MediaEval grows stronger by staying in close contact with groups running benchmarking activities in areas beyond the MediaEval core domain of human and social multimedia. MediaEval 2011 was presented at CLEF 2011, NTCIR 2011 and FIRE 2011 and in 2012 we hope to convince some of our sister benchmarks to give a reciprocal presentation on their own experiences.

As part of staying connected to the larger community, the MediaEval 2011 workshop included a poster and demo session where projects that help to organize MediaEval tasks could present their results and where industry people could make a presentation of new and interesting problems that they would like to make known to the MediaEval community as possible benchmarking tasks.

The workshop closes with a gathering of task organizers and others who have invested time and effort into the community or want to get more closely involved in the future. During this session, we reflect back on what happened so far in the benchmarking year and also discuss the MediaEval related activities that we organize beyond the core benchmarking activities. These include joint papers among all the participants of a task and also special sessions at conferences. Looking forward we also plan to get involved in organizing more workshops (of the traditional variety) at conferences and also think about the possibilities of special issues. Finally, we reflect on our MediaEval goal, To offer the community innovative new tasks related to the human and social aspects of multimedia and our slogan The "multi" in multimedia (it not necessary to be able to say the slogan with a straight face.) On the basis of these reflections, we consider where we would like to go with the benchmark in the coming year.

So why do you want to be at the MediaEval workshop? Well, if you are a task participant, it gives you an opportunity to exchange with other people working on topics similar to yours and helps you to understand and improve the algorithms that you are using to approach your task. MediaEval needs a central group of dedicated researchers to organize the tasks that make the benchmark run. Attending the workshop is a good first step to getting more deeply involved in MediaEval, for example, by proposing a task for the year.

On the MediaEval 2012 survey, we had a question concerning what the community thinks about how MediaEval should grow. I personally, want to keep the workshop as small and intimate as possible. In 2011, we had nearly 60 people and that appeared to me to be a good maximum size for the workshop. However, when I examined the survey results, I realize that I am in the minority here, and that most of the people in the community would like growth. As a result, I am changing my opinion and we will not attempt to artificially restrict growth, as long as it is sustainable.

The issues and ideas around growth is just one area in which the concept of the MediaEval workshop is evolving. One of MediaEval's strengths is that it develops from year to year, guided by the input of the community --- and in particular those people who invest the most hours of their time to make it work. I look forward witnessing and being involved in this development in the 2012 season and, we hope, to seasons beyond.

For further impressions of the MediaEval workshop, check out the MediaEval 2011 workshop video: