Showing posts with label concepts. Show all posts
Showing posts with label concepts. Show all posts

Friday, December 12, 2014

"Smart Photography" is not So Smart

This post enumerates the reasons for which "Smart Photography" technology described in this article: 
Stokman, Harro. The Future of Smart Photography. Computing Now. IEEE Computer Society. pp. 66-70, July-September 2014
is an enormously bad idea. The technology attempts to categorize images at the moment at which they were taken, and prevent certain types of images from being taken in the first place. In this post, I point out the reasons for which no analogous technology exists for text production. Then I go on to argue that "Smart Photography" constrains the ability of people to record important moments, and, critically, could hinder the ability of a witness to collect evidence during a crime. I close with an example of an image, with which I exercise my freedom of expression. 

I am writing a blog post. Let's think for a moment about what is not happening. Specifically, let's reflect on the fact that www.blogger.com is not immediately attempting to put this post into a particular topic category as I write.

One reason why it is not attempting this is that, ultimately, automatic prediction of the topic of my writing is not particularly useful to me. Automatic topic detection could misinterpret my topic, or it could completely miss the fact that I was writing on a new topic that had never been written about before. My post would get misfiled, and potentially lost.

Such a text classification technology would also need to assume that topic was indeed the appropriate type of category into which I wished to sort this blogpost. There would be no room to invent new types of categories, related to e.g., style, place, sentiment or mood.

From this example, we observe: Blogging is a couple decades old, arguably older. After all of this time, www.blogger.com keeps the responsibility of categorizing blogposts firmly with the writer. I will need to click the "Labels" icon and add the tags myself.

Here is another example: As I go through a series of Mac computers, they have gotten more sophisticated over the years. However, there is no functional "autosuggest" feature that can predict in which folder I will want to store a document or presentation while I am creating it. Because of the nature of human creativity, it just doesn't make sense to do this if the computer is going to be useful for tasks that have not yet been imagined.

Finally: YouTube auto suggests categories (presumably using my metadata), with the effect that all my videos are "Science & Technology". That helps me to find, well, exactly nothing in my uploads. Everything is labeled the same. There, again, it's clear that creation must also involve classification effort, if the end effect is to be organization. We don't create towards a pre-defined set of concepts or topics. Instead, when we create content, we also create concepts.

These three examples illustrate that the idea of classification of text at the moment of creation, has not "caught on". And, it's not because we do not yet know how to train text classifiers. Text classification technology has long been considered far ahead of computer vision. We need to acknowledge that other forces are at play.

Yet, in the face of a lack of general applications that classify text at the moment of creation, computer vision researchers are now attempting to build "Smart Photography" applications that would classify images at the moment they are taken by cameras. This contradiction should make us really sit up and think hard about the implications of "Smart Photography".

"Smart Photography" is first fascinating, and then horrifying.

It's fascinating because it's fun. If my camera decided at the moment I clicked the shutter that I was taking a picture of food, I probably would take more pictures of food. I would do it because it would be cool to see if the camera "understands" food. (Yes, I'm a multimedia geek). Also: because food as a separate category is actually built into my camera, I would feel less embarrassed about taking out my device and snapping a picture before eating.

However, this kind of behavior effectively amounts to the camera teaching me what I should be taking photos of. It will subtly channel human photographic impulses down the broad and easy road, and allow less traveled paths of expression to slowly grow over.

That's sad, but not yet horrifying. Horrifying is the following:

The "Smart Photography" camera is going to prevent its user form taking certain photos entirely.

Reading the "Smart Photography" article, the computer vision experts obviously have their hearts in the right places when they envision a camera whose shutter freezes when part of a hand appears in the frame.

However: Imagine your baby's first steps. You miss the shot because you just couldn't get your finger out of the frame in time. You would much rather have a "baby plus finger" photo and be able to treasure the moment, than have your camera freeze up on you because it "saw" a finger in front of the lens and locked the shutter.

It goes on: You miss the shot of your kid scoring that amazing soccer goal, of that rare bird that you saw on your walk in the woods. You miss the shot of the damage that was done to your car in an accident because you just couldn't hold the camera perfectly correctly. A bad shot would have been better than none.

And it gets worse. The camera aspires to block adult content, making it impossible to take a picture of any scene that it classifies as pornographic. It sounds like a miracle for law enforcement the first time you hear it. But the price is too high: Basically, if the camera blocks adult content, it means that if I witness a rape, I have no way of taking a picture of it. The possibility of identifying the perpetrator is blocked by the camera itself.

Effectively, the innovation of "Smart Photography" is making possible a camera that does not work.

Returning to the comparison to the text case. What if www.blogger.com was preventing me from writing this column, as soon as it sensed that its topic included rape? Our technology does not prevent the generation of text, and we need to remain consistent with the values that tell us that lead us to the conclusion that it should not. It is a bad idea to introduce technologies that prevent the generation of images.

One important reason why our technology does not censor text on creation, is it is not people who design technology that get to make the decision about what I can and cannot express. Rather it is the legal system, which is in turn based on the values of the community at large. This system, imperfect and slow moving as it might be, represents individual citizens equally, and can be influenced by them, in the way that a technology cannot be influenced equally by everyone.

The "Smart Photography" article argues that photocopiers prevent people from copying money, i.e., paper bills, and that this technology represents a next step. The reason why money works at all is that there is a system and a society working to make sure that its purpose is unambiguously interpreted. Money is a conventionalized sign at the basis of the society, it exists at all exactly because it is not open for interpretation. Plainly stated, the argument that technology blocking photographic capture of adult content is a natural extension of blocking photocopies of money, relies on an unsound analogy, and must be discounted for this reason.

What's the alternative to "Smart Photography"?

The solution is not making photos smarter, but rather it's changing people. It's relentlessly pursuing our efforts to support each other in our communities, and to help each other make better decisions. It's about the unending quest, that begins again with each new generation: to make people smarter.

We need to assure adequate funding to the people who dedicate their careers to fighting crime. Finding perpetrators of sexual abuse/sex crimes is simply a hard task that requires a huge investment: sick minds are sick, and they will not let a new camera technology stand in the way of their evil business. With this "Smart Photography" camera, sex offenders will be incentivized to start taking pictures that are not so easy to automatically identify, and they may be able to wipe out their own footprints. There are no easy technical shortcuts that will eliminate the need for old fashion crime fighting, yes, also of the gumshoe variety.

We need to educate people. Bear selfies are stupid. Getting people to stop taking bear selfies is not a matter of creating a camera that recognizes a bear selfie situation, and blocks the shutter when someone tries to take a bear selfie. The bear selfie is a symptom of an underlying lack of reflection. It is the underlying problem, and not its superficial manifestation that needs to be addressed. The answer is about taking the time to really talk to our children, and to each other, about what is appropriate and what is not appropriate in a given situation.

Below is the most repugnant photo that I have ever posted online, but today for the first time in my life, I did not take for granted that my camera includes a functionality that allows me to take it.

Thursday, May 17, 2012

Search by misconception: Should search engines support information needs that are ill conceived?


The real Mozartkugeln? (Flickr: davidroethler)
Should our search engines make information findable based on misconceptions? Well, it's complicated. This post gives some examples to highlight the relationship of misconceptions to search.

Let's start with an example. When I refer to "Mozartkugeln" I mean the ones in the pictured here. They are gold and show Mozart in his red jacket. Their shape is round. I differentiate these from the ones with flat bottoms, which are for me "fake" Mozartkugeln.

My idea of Mozartkugeln can be considered a misconception. The original Mozartkugeln are apparently produced by a company called "Fürst" and are silver with a blue Mozart. Additionally, the producer of the flat-bottom ones apparently has the right to call their product "Real Reber Mozartkugeln". Digging on Wikipedia and on other websites supplied me with this information.

But what should an image search engine return in response to the query "Mozartkugeln"? Is it obliged to make an effort to resolve the question of which is the "real" Mozartkugel? Or is it fine if it just returns images that users have uploaded a tagged with "Mozartkugel"?

Effectively, simply returning images tagged with "Mozartkugel" allows users to search by misconception. The search engine returns images who have been tagged by people like me, who have a certain view on the matter (based on conversations with an Austrian roommate now a couple decades old and several subsequent trips to Austria, none including Salzburg), which is not necessarily universal. I am not immediately convinced that I can be satisfied with such a search engine as a source of information. Although, it seems reasonable to assume that if enough voices are combined, a consensus will emerge. I noticed that if you search for "champagne" on Google images, the top hits (at least the ones that depict identifiable bottles) clearly hail from the Champagne region in France and don't include the large range of other bubbly wines from other corners of the world that are widely enjoyed under the name "champagne".

In short, allowing search by misconception seems relatively innocuous. But we should be careful about assuming that the Mozartkugeln example is the end of the story. What is unique about this example, is that the search engine is relatively transparent in the way that it works. The images are returned by seeking exact matches in their user-assigned tagsets; without such a match, the image is not relevant. Users of the search engine have a chance of being at least vaguely aware of the reason for the match and they can propagate their understanding of the reliability of the taggers to create an understanding about the reliability of the results.

However, when the search engine becomes more sophisticated, the situation quickly gets quite murky. For example, if I had a visual concept detector that was trained to detect Mozartkugeln in images and assign to them the appropriate tags. The design of the detector would require collecting examples of Mozartkugeln, which means that whoever trains the detector holds the ultimate control over deciding what a Mozartkugel actually is.

The example of Mozartkugeln is interesting. In some cases, one could argue that common sense knowledge will tell you what an object is, for example, a helicopter or a pram. Everyone can identify these objects, right? But in the case of the Mozartkugeln, there is no right answer. It depends on your perspective. A long discussion will arrive at the conclusion "It's complicated". (And you may already find yourself with the same issue for the pram, if not actually the helicopter.)

It seems like a good idea to do away with the central authority that collects the examples used to train the detectors. After all, no one really likes they guy who walks around the party reminding people, "Yes, but it's not real champagne".

But do we really want to admit search by misconception? I had quite an unsettling experience with Google's query suggestion. On 17 May 2011, I was looking for a news story on one of the Facebook founders have renounced his US citizenship. No sooner had I typed in "facebook founder" did Google present me with the following list of suggested queries:

facebook founder mark
facebook founder saverin
facebook founder buys new republic
facebook founder college
facebook founder gay
facebook founder bios
facebook founder dead
facebook founder movie

Did I really need to know about the existence of the circulating rumors? Do I go on to passively "believe" a query or do I dig deeper that find out if it is true? Do we really want our search engines to allow us to so easily flow down the same information paths worn by searchers before us who mis-received a rumor?

In the case of "facebook founder dead" I did dig deeper. That query led to a Fox News article on the death of Ilya Zhitomirskiy, one of the co-founders of Diaspora*, an alternative to Facebook. I was left wondering at how query suggestions have taken on an information dissemination (news broadcast, if you will) role of their own.

From the fun of searching for pictures of bonbons (...and wondering if round vs. flatbottom Mozartkugel relates to a real misconception or rather an alternate interpretation) we hit on a matter of true importance (Diaspora* upends the Facebook model because it is based on the idea that every member of the network should “own” his personal information). Suddenly, it gets extremely serious. In light of this seriousness, it looks like if we really do not want search engines to admit search by misconception at all.

This whole line of though was started while I was at a symposium entitled "Cultural Heritage Gets Social" of the SEALINCMedia project (Socially-enriched access to linked cultural media). Alice Warley (Public Catalogue Foundation, UK) gave a talk entitled "Your Paintings Tagger: Crowd-sourcing, art history and the UK's national oil painting collection" about a website where visitors collaborate to tag painters

Apparently, general public users tend to tag older paintings "formal wear" when what the people pictured in the paintings are wearing is not formal wear at all, but rather daily clothing.
The reason for this misconception is that today's formal wear evolved from what was worn on a daily basis in certain social circles in past eras. Wikipedia is rather silent on the history of formal wear. The "misconceptions" of the taggers are actually a source of information about something that is not widely known, but actually an real historical connection.

So we're back to the Mozartkugeln, considering whether Mozart is dressed for a concert, or is in his everyday work clothing that he uses for composing. It seems like misconceptions help us to uncover new and interesting information.

However, if we incorporate misconceptions, maybe we should call them 'exploration engines'. A 'search' engine should find answers or else gently reveal to us that our initial information need was ill conceived.

Monday, October 31, 2011

Reflections on visual concepts in images (Halloween I)



LSCOM stands for "Large Scale Concept Ontology for Multimedia" and it is a list of concepts associated with multimedia, including images and videos. If you are to ask me where I stand with the LSCOM concept list, I am a 2753-Solid_Tangible_Thing kind of a multimedia researcher and not a 125-Airplane_Flying kind of a multimedia researcher.

Basically, what I mean is that I adhere to the perspective that in order to solve the general problem of multimedia information retrieval on the Web, we should make use of basic properties of objects depicted in images and video, rather than their specific identities. I have discussed the issue previously in a post on proto-semantics, dimensions of meaning that arise from human perceptions and interactions with the world. Proto-semantic dimensions are more fundamental than the words that we usually use to describe the world around us, and for that reason, they can be considered to be sub-lexical. For example, I am drinking coffee from a mug, but more fundamentally this is a small, corporeal object, or if we pick something from LSCOM 1425-Concave_Tangible_Object. I return to the issue here, since I've been pondering it again on the occasion of Halloween.

It seems that the way that scientists approach the problem of visual indexing, i.e., automatically describing the visual content of images and videos, is always inextricably related to their backgrounds. I've worked in the area of multimedia retrieval for going on 12 years now, and it my experience two main backgrounds dominate the field: surveillance and cultural heritage. Let me say a few words about both.

Surveillance: The analysis of surveillance footage or images captured by security cameria is aimed at the task of automatically identifying threat levels. For surveillance tasks, one defines a closed set of objects and behaviors that constitute "business as usual" and anything outside of that range can be considered a threat and triggers and alarm calling for the intervention of human intelligence. Surveillance is a high recall task -- meaning that it is more important not to miss any events than to reduce the detection rate of false alarms. This background doesn't quite transfer to the general problem of multimedia retrieval on the Web.

We can't assume that Web multimedia will depict a closed class of objects. The cases that cannot be covered by a closed class are not infrequently occurring "threats", but rather entities drawn from the long tail: which, if we can indeed assume a finite inventory, will contain approximately half of the encountered entities. Further, Web multimedia retrieval is typically a precision oriented problem, which means that reducing false alarms is relatively more important than exhaustive detection.

Cultural heritage: Iconographic classification of visual art involves a classification system such as Iconclass. The stated purpose of Iconclass is the description and retrieval of subjects represented in image. I rather suspect that before the very first paint had dried on the very first canvas, next to the artist was standing an art historian who started to create a classification system to categorize the painting. In other words, using classification systems for visual art is an old idea, that has well-established conventions and has been honed over generations of use. Such a classification system necessarily views works of art as physical objects, and would have as it's goal the task of organizing the storage facility of a museum or of helping to choose which works to hang together in an exhibition. The people who created it assumed that the number of dimensions of similarity between works of art was necessarily finite. Such an assumption makes sense, in light of a relatively small number of art historians working on a relatively small number of questions concerning art history and the iconography of art.

Enter, however, the Web. Images and video are not physical objects and we do not have to be able to list them all in a well ordered list or even every make the decision of "Do we hang this in the East Wing gallery or the West Wing gallery?" There are many more users than art historians, and suddenly it actually be useful to admit the possibility that the number of ways to compare two images might in fact be infinite, rather than finite.

As for myself, I neither fall into the surveillance or the cultural heritage category. I attribute this to what's probably a naive equation of surveillance with totalitarian states and also to having the yearly experience in grade school of being packed on a bus and shipped off for a day at the Art Institute of Chicago.

I guess the Art Institute of Chicago was supposed to have broadened the horizons of our young minds, but instead it sort of warped me in a way that makes it difficult to talk to me, if you are an cultural heritage person or an art historian. I was young enough that everything I drew sort of came out flattish, whether I intended it to look two-dimensional or not, when I was suddenly confronted with the likes of Marc Rothko. I think what happened is that someone in Chicago told me that Marc Rothko described his work as an “elimination of all obstacles between the painter and the idea, between the idea and the observer” (as quoted on this AIC webpage describing the Rothko painting above). At the time, I didn't particularly like Rothko, but the experience permanently hardened my mind to the idea that it made any sense whatsoever to describe visual art in terms of its depicted subject.

I think that Marc Rothko must fit into iconclass categorization "0 Abstract, Non-representational Art: 22C4 colours, pigments, and paint", which is unsatisfactory to me because it makes him seem like an afterthought. In Chicago, they apparently forgot to mention that he was reacting to what came before him. For me, I was already broken. A system that put Rothko on the outside rather than at its core could never been acceptable to me. From then until always: the main point of art is what we do with it: how we talk about it, how we stand before it and mull in the museum, which prints we buy in the shop and go home and hang on our walls and (as little as we like to admit it) how much we pay for it. A priori we don't know what draws us to art, so why should we make little lists of entities corresponding to its subjects?

The perspective I take may not ultimately prove more productive than either the surveillance perspective or the cultural heritage perspective. It is the linguistics perspective. My view is the following: the elements of meaning arising from human perception and interaction with the world that have been encoded into language human language semantics, these are the elements that we should try to dig out of videos and images. They are the lowest common denominator of meaning that we can be sure will give us the ability to cover all human queries: the ones that we can anticipate and the ones that we cannot.

So should the image above be given the LSCOM category 2753-Solid_Tangible_Thing ? Sure. It's an image of a painting. That's a tangible object. But let's also let the image be found by shape and color. And be found how I found it on the Internet: with the query "Rothko". And let it also be found when we search for formative experiences. And for Chicago...

And what does this have to do with Halloween...continue to the next post.