Showing posts with label Data visualization. Show all posts
Showing posts with label Data visualization. Show all posts

Saturday, March 5, 2011

Tips for data mining - Part 2 out of many

Writing after a long time on the blog. Blame it on regular writer's cramp - a marked reluctance and inertia of sorts to put pen to paper, or rather fingers to keyboard.

My last post introduced the idea of defining the problem as the first step for any data mining exercise aiming to achieve success. This is the ability to state the problem you are trying to solve in terms of business outcomes that are measurable. After that comes the step of envisioning the solution and expressing it in a really simple form. The aim should be to create a path from input to output - the output being a set of decisions that will ultimately result in the measurable business outcomes we mentioned above. The next step involves establishing how the developed solution would be used in the real world. Not doing this early enough or with enough clarity could result in the creation of a library curio. Defining how the solution will be used will also point to other needs such as training the users on the right way to use the solution, the expected skills from the end users and so on.

In this post, we will discuss the third and fourth steps. These are
3. Frame the approach before jumping to the actual technical solution
4. Understand the data

Frame the approach before jumping to the actual technical solution. Once the business problem has been defined, it is tempting to point the closest tool at hand at the data and starting to hack away. But often times, the most obvious answer is not necessarily the right answer. It is valuable to construct the nuts and bolts of the approach to get to the solution on a whiteboard or sheet of paper before getting started. Taking the example of some recent text-mining work I have been involved in, one of the important steps was to create an industry-specific lexicon or dictionary. While creating a comprehensive dictionary is often tedious and dull work, this step is an important building block for any data mining effort and hence deserves the upfront attention. We couldn't have seen the value of this step, but for the exercise of comprehensively thinking through the solution. This is also the place where prototyping using sandbox tools like Excel or JMP (the "lite" statistical software from the SAS stable) becomes extremely valuable. Framing this approach in detail allows the data miner to budget for all the small steps along the way that are critical for a successful solution. It also enables putting something tangible in front of decision makers and stakeholders which can be invaluable in getting their buy-in and sponsorship for the solution.

Understand the data. This is such an obvious step that it has almost become a cliche; having said that, incomplete understanding of the data continues to be the reason why the greatest number of data mining projects falter in attempting to fulfill their potential and solve the business goal. Some of the data checks like data distributions, variable attributes like mean, standard deviations, missing rates are quite obvious but I want to call out a couple of critical steps here that might be somewhat non-obvious. The first is to focus extensively on data visualization or exploratory data analysis. In the blog, I have written a few pieces before on data visualization which can be found here. Another good example of this type of visualization is from the Junk Charts blog. The second is to track data lineage - in other words, where did the data come from and how was it gathered. Also is it going to gathered in the same way going forward. This step is important in understanding whether there have been biases in the historical data. There could be coverage bias or responder bias, where people are invited or requested to provide information. In both these cases, the analytical reads are usually specific to the data collected and cannot be easily extrapolated to non-responders or people outside the coverage of the historical data.

This covers the background work that needs to take place before the solution build can be taken up in earnest. In the next few posts, I will share some thoughts on the things to keep in mind while building out the actual data mining solution.

Thursday, January 13, 2011

The Joy of Stats is finally online

My first post of 2011. I have been writing this blog for nearly two years now and am happy to keep having the energy and the enthusiasm to keep at it. Like I have said earlier on why I blog, this is way for me to keep abreast of the latest development in the fields of data mining, analytics and visualization.

The Joy of Stats program was aired in BBC4 in December 2010. Now the video is available on Hans Rosling's Gapminder website. This was the program from which the data visualization examples used for mapping San Francisco crimes, the graphics made by Florence Nightingale and the Gapminder visualization of the economic and demographic statistics of different countries over the last 200 years are highlighted.

Another example of nifty graphics. The New York Times has been a trendsetter in putting up very clever and informative graphics supporting its news stories. Amanda Cox of the NYT graphics department did a presentation recently on some of the examples that the NYT has used in its print as well as its online media. This is a long presentation but worth sitting through.

Hopefully you will enjoy both these presentations!

Saturday, December 18, 2010

Visualization of the data and animation - part II

I had written a piece earlier about Hans Rosling's animation of country-level data using the Gapminder tool. Here are some more examples of some extremely cool examples of data animation.

At the start of this series, there is more animation from the Joy Of Stats program that Rosling hosted in the BBC. The landing page is a link that shows the plotting of crime data in downtown San Francisco and how this visual overlay on the city topography provides some valuable insights on where one might expect to find crime. This is a valuable tool for police departments (to try and prevent crime that is local to an area and has some element of predictability), residents (to research neighbourhoods before they buy property, for example) and tourists (who might want to doublecheck a part of the city before deciding on a really attractive Priceline.com hotel deal). The researchers who have created this tool that maps the crime data to maps. The researchers in the clip talk about how tools such as this can be used to improve citizen power and government accountability. Another good example of crime data, this time reported by Police Departments across the US can be found here. Finally, towards the end of the clip, the researchers go on to mention what could be the Holy Grail of this kind of visualization. They talk about how real-time data put up on social media and networking sites like Facebook and Twitter (geo-tagged perhaps) could provide a real-time feed into these maps. Now this would have been certainly in the realm of science fiction only a few years back but suddenly now it doesn't seem as impossible.

The San Francisco crime mapping link has a few other really impressive videos as you scroll further down. I really like the one of Florence Nightingale, whose graphs during the Crimean war helped reveal important insights on how injuries and deaths were occurring in hospitals. It is interesting to know that Lady of the Lantern was not just renowned for tending for the sick, but also was a keen student of statistics. Her graphs of deaths which were accidental, caused by war injuries and wounds and finally those that were preventable (and caused by poor hygiene that was quite prevalent at the time) created a very powerful imagery of the high incidence of preventable deaths and the need to address this area with the right focus.

Why is visualization and animation of data helpful and such a critical tool in the arsenal of any serious data scientist? For a few reasons.

For one, it helps tell a story way better than equations or tables of data do. That is so essential to convey the message to people who are not necessarily experts who have insight into the tables, but are important influencers and stakeholders nevertheless who need to be educated on the subject being conveyed. Think of it as how an advertisement (either picture or moving image) is more powerful in conveying the strength of a brand as compared to boring old text.
The other reason, in my opinion, is that graphical depiction and visualization of the data allows the powerful human brain (which is far more powerful than any computer at pattern recognition) to take over the part of data analysis that the human brain is really good at and computers generally not so good at. This is forming hypotheses on-the-fly about the data being displayed and reaching conclusions based on visual patterns in the data. Also the ability to hook into remote memory banks within our brains and form linkages. While Machine Learning and AI are admirable goals, there is still some way to go before computers can match the sheer ingenuity and flexibility of thought that the human brain possesses.

Tuesday, November 30, 2010

Animating the data and better "story telling"

One of the challenges with talking about and presenting any analysis about data mining or statistics is that a lay audience is seldom excited by the same things as a more technical audience. A technical audience is as interested in how the answer was reached as much as the answer itself. A non-technical consumer of the same information is probably interested in the implications of the answer as well as the answer itself, with some gut-check to make sure that the process wasn't totally crazy. In other words, they are looking for a story.

Recent trends around the pervasiveness of data and data-driven applications has meant that there is a greater ask from data scientists to tell a compelling "story" to support their analysis. Data scientists need to come up with ways that tell the story behind the data and the projections of the model that may have used the data as input, that are insight generating, that skip some of the unnecessary detail and also paint the various facets of the final solution. And not just tell the final answer. Data animation and data visualization are some of the answers here.

I came across a couple of good examples of such animation recently. Hans Rosling made an interesting presentation about a tool called Gapminder at TED.com. The presentation is here. Gapminder is an organization that makes social, environmental and economic development data from all the countries of the world available and accessible to all, for free. The visualization tool at Gapminder called Gapminder World shows ways in which this data can be animated and made come alive for the non-technical consumer in illuminating and exciting ways.

Rosling made another trailer presentation recently from a BBC 4 program promo called "The Joy Of Stats". (Link is here if the embed doesn't work.). One hopes that this program airs sometime in the US. It is due to air on Dec 7 and 8 in the UK. Any UK readers of the blog are encouraged to go and check the program and share what they felt about it. The content of the program (Link and timings here) sounds interesting enough for me to at least contemplate taking a flight to London and catching the program on the Beeb.

Tuesday, May 4, 2010

Interesting data mining links

1. The NY Times recently had a piece on how data is increasingly part of our life. Link here.

2. The Web Coupon - a new way for retailers to know more about you. Link here.

3. On Principal Components Analysis. Link here.

Tuesday, July 21, 2009

More data visualization - this time about books

Ever wonder where the proof was about reading ..ummm, erotica being bad for you. Here it is. Check this link out.

An interesting study was done that went somewhat like this.
- Get the ten most frequent "favorite books" at every college using the college's Network Statistics page on Facebook. Possibly these books represent the intellectual calibre of the college.
- Get their SAT/ACT scores for the colleges.
- You can now get a relationship between types of book read and scholastic achievement

The results are pretty impressive, though still somewhat dubious. According to the study, Classics is usually good for you (agree with that), Erotica is way bad. Controversially, so is African-American literature and chick-lit. In the link, check out the visual that stacks the book by genre.

Make what you want about this, but be careful between causality and correlation.

Saturday, July 18, 2009

Data visualization


An example of a really well-done graphic is from the NOAA website. Science and particularly math afficionados seem to have a particular affinity to following weather science. (I am wondering whether it is a visceral reaction to global warming naysayers who, the scientists think, are possibly insulting their learning.)

The graph is a world temperature graph and this type of graph has come in so many different forms, it is difficult not to have seen such a graph. What I like about this is the elegant and non-intrusive form in which the overlays are done.
• By using dots and varying the size of the dots, the creator of the graph is making sure that the underlying geographic details (important in a world map where there is great detail that needs to be captured in a small area, therefore you cannot use very thick lines for country borders) still come through.
• The other thing that I liked is some of the simplications the creator has made. The dots are equally spaced but I am pretty sure that’s exactly not how the data was gathered. But to tell the story, that detail is not as important.

The graphic came from Jeff Masters' weather blog which is one of the best of its kind. Here's a link if you are interested.

Friday, May 29, 2009

Data handling - the heart of good analysis - Part 1

I have been building consumer behaviour models using regression and classification tree techniques for the last 4 years now. Most of this work has been in SAS. Now, there are a large number of interesting SAS procedures that are only slightly different from one another. Many of them can be interchangeable used, like PROC REG and PROC GLM.

But the single most important learning for me over this period has been that you can't spend enough time understanding and transforming the data. Very many interesting and potentially promising pieces of analysis go nowhere because the researcher has not enough time understanding the data. And then, having understood the data, transformed it into a form that is relevant to the problem at hand.

One of the seminal pieces on understanding data and plotting it in useful ways, is John Tukey's "Exploratory Data Analysis". This paper introduces some unique and important ways of graphing and understanding what the data is trying to say. One of my personal favorite SAS procedures is PROC MEANS and PROC UNIVARIATE. And of course PROC GPLOT. My advice to the budding social scientist and quantitative practitioner is to learn to use these techniques before learning the cooler procedures like LOGISTIC and Linear Models. This was one of the first things I learnt in my own journey as a statistical modeler and I have some very good and experienced colleagues to thank for making sure I learnt the basics first.

Over the next several posts, I am going to share some of my favorite forms of data depiction. The next several days will be a very interesting read, I promise.

Sitemeter