Showing posts with label self-organizing maps. Show all posts
Showing posts with label self-organizing maps. Show all posts

Friday, January 05, 2007

The dirt on Ting-lan

Ting-lan was thoughtful.

She had solved a problem sideways the way mathematicians go at it sometimes.

The problem had been to match answers that had no questions to questions that had no answers.

Ting-lan smiled. Papa had been no good. If Mama could see her now. She was about to wonder twice but decided this was not the time for Ting-lan Tao. By Ting-lan Tao Ting-lan meant the reverie.

Ting-lan's solution came from Dirt. Dirt is an online celebrity gossip magazine where fans can submit audio comments to podcasts and spill their beans. Not mung beans but all kinds of confession beans. Tittle beans. Tattle beans. Audio beans.



Dirt is powered by Big In Japan. Big In Japan has an ethos for developers. Ting-lan followed the ethos.

In this ethos Ting-lan became a "social samurai". A social samurai would not dream of clustering answers and questions through their feature values without user input.

In Ting-lan's solution a user/interviewer enters the bestmatch algorithm via The Door and becomes part of the program. Ting-lan built The Door using parts from Big In Japan.

There was a movie once where users entered programs -- Tron. Ting-lan's solution was Tron for interviewers.

Who said survey research wasn't cool? Cooler than Google? Yes, Ting-lan thought, Google is boring.

Read More...

Wednesday, January 03, 2007

Google is boring

Google is boring. Quintura is not.

I enter "self-organizing map". Along with the usual result set, I get a cloud. In the cloud I see my keyword along with other ones sized according to their frequency.




If I want to add a keyword to the cloud that is not already there, I just double click:




Now my result set is the product of the two keywords. Pretty cool! It gets better. I hover on "cluster analysis" and preview its keywords:



Finally, I select "cluster analysis". And I get a new cloud and a new result set. They are the products of "self-organizing map" and "cluster analysis":




Not bad. Better than not bad. Not boring...

In a recent post at Read/WriteWeb Alex Iskold wrote otherwise at The Race to Beat Google. He was somewhat dismissive of the clustering solutions saying they were too complex to make it in the mainstream. I don't think we want to underestimate clustering when it is wedded with clouds. Indeed clouds informed by clustering like the ones above can rain on Google's parade. I hear from kids that they always wanted clouds to do the walking. Imagine a swim in clouds that have three dimensions. Instead of "googling", the euphemism will be "I need to check the weather." The rejoinder from Will Smith's wife in Enemy of the State would be: "And just whose weather are you checking?"

Read More...

Tuesday, January 02, 2007

Velvet has a tangerine day

Velvet was having a tangerine day. On a tangerine day every interview she touched would flow and glow and she was reminded of that song by Talking Heads which right now she couldn't stop hearing even if she wanted to:

I dont know why I love you like I do
All the troubles you put me through
Sixteen candles there on my wall
And here am I the biggest fool of them all

I wanna know that youll tell me
I love to stay
Take me to the river and drop me in the water
Dip me in the river, drop me in the water
Washing me down, washing me down.

Yep. It was one of those drop me in the water days.

There had been alot of those days since the retooling. The retooling was when she sent her laptop back to the company. Then, a few days later she got back a Wii.

It wasn't really a Nintendo. It was still a laptop. But it was also a Wii.

It was a Wii because now when she dragged the external data sources avatar onto the interview and there were still too many unanswered questions, she could shake and bake.

Shake and bake was a new resource. It used something called "connecting the dots" to turn unanswered questions into answered ones. The reason they called it "shake and bake" instead of "connecting the dots" was easy. Because you just didn't activate the shake and bake plan by dropping it on the interview and its lost sections. After you dropped the plan, it had to be simmered and cooked. That was when Wii came in.

Velvet would pick up her laptop and move it from side to side and back and forth until there was a fit and the answers covered all the questions.

The company never really explained the Wii except to say that the laptop was now capable of doing "lateral thinking" but the thinking had to be shaped. Hence shake and bake and Wii. And hence visualization. And hence the interviewer as cook.

It was all very natural. Smile.

Velvet wanted to look under the cover of shake and bake just like she would look under the cover at her boyfriend. But she knew this time she wouldn't understand anything she saw. Even so, Velvet couldn't let sleeping dogs lie. The sleeping dog in this case was a tip -- you know, one of those yellow popups that surfaced under a mouseover.

When Velvet moused over the shake and bake plan she got the tip. The tip said "self-organizing map". Sometimes it simply said "SOM". And sometimes there was the video.

Yep. Take me to the river and drop me in the water. It wasn't better than sex but Wii was it fun.

Read More...

Monday, January 01, 2007

Technorati Charts

Posts that contain "data mining" per day for the last 30 days.
Technorati Chart

Posts that contain "cluster analysis" per day for the last 30 days.
Technorati Chart

Posts that contain imputation per day for the last 30 days.
Technorati Chart

Posts that contain "self organizing map" per day for the last 30 days.
Technorati Chart

Read More...

Friday, December 29, 2006

An Interlude: Self-Organizing Maps

On September 28, 2006 the US Patent Office published a patent application from Microsoft entitled System and method for improving search relevance.

The inventor contemplates the following problem:

Take a collection of documents, say about the size of the Web, and try to organize them based upon textual similarities between them. Can that organization provide a useful way to index the web?
The invention would augment keyword search. Documents would not be indexed based on keywords directly. Instead there would be an indirection. Documents have labels -- many labels. Just look at how documents are labeled by a typical user of del.icio.us or one of its competitors. Keywords would be related to labels that are related to documents.

Some invention like this is what Microsoft proposes using a technique called self-organizing maps.
The Self-Organizing Map (SOM) by Kohonen is motivated by the receptive fields in the human brain. High dimensional data [e.g. labeled documents where each label is a dimension] are projected in a self organizing process onto a low dimensional grid [e.g. a system of keywords that Microsoft refers to as "content tiles" in the application] analogous to sensory input in a part of the brain.
See the discussion of Emergent SOM at the website of the Databionics Research Group for a more in depth treatment of self-organizing maps including some nifty visualizations of the SOM process. See also my del.icio.us som.

Meanwhile here is the patent application abstract:
A system and method for performing context based document searching is provided. A grid of content tiles is constructed corresponding to a desired concept space. Each content tile is assigned a content tag and is associated with a series of feature values. The feature values are trained to correspond to various regions of the content space. Documents are associated with one or more content tags based on a comparison of document feature values with content tile feature values. A search query is modified to include one or more content tags based on the terms in the search query and/or user preferences. The search query is then matched to documents associated with content tags contained in the search query.

Read More...