Sunday, August 9, 2009

Knowledge-worker roles in the 21st century - 1/2

It is rare to find a piece in the media nowadays that doesn't have a certain "view" on the important social and economic issues of today. Underlying every opinion piece is ideology of some sort. Slavish commitment to the ideology results in the writer typically producing such a biased view that it only appeals to those who are already prejudiced with the same view. It has become extremely difficult (especially with the evolution of the Internet and with blogs) to reach an informed and balanced view on a subject by referencing an authoritative piece on the subject.

Which is why I was pleasantly surprised to come across this piece in the Washington Post today. The piece by Gregory Clark, a professor of economics at UC, Davis presents the view that many of us are afraid to admit. And that is that the US will very soon be forced to confront a reality where the technological advances in the economy today creates its own Haves and Have-Nots. And the chasm between the Haves and Have-Nots would be so huge and so impossible to bridge that the government will be forced to play an equalizing role, so that the social order in society remains more or less intact. So how is new technology creating this chasm? More importantly, for me and my readers, what are the kinds of knowledge-worker jobs that are going to be valued in the twenty-first century?

The last fifty years of the second millenium have been marked by the emergence of the computer. A machine designed to do millions of logical and mathematical operations in a fraction of a second, the computer has now started to take over a vast majority of the computing and logical thinking that human beings would usually perform. With the ability of the computer (through programming languages) to execute long sequences of operations at high speed, the end-result is a powerful "proxy" intelligence that can be harnessed to do both good and harm. And this proxy-intelligence is taking the place of traditional intelligence; the role performed by human beings in society. And this intelligence comes without moods, expectations of recognition/ praise; in fact, without any kind of the emotional inconsistencies and quirks shown by human beings. No surprise that many of the front-end business processes involving the delivery of basic and transactional services to consumers is being replaced by the computer (such as the ATM machine). With the computer becoming an increasingly integral part of the economy, I see two kinds of jobs that knowledge-workers can embrace in this economy. I am going to cover one of these roles in this post and the second, in the next post.

The first role is that of the accelarator towards an increasing automation of simple business processes. The cost benefit of the computer over human beings is obvious; however in order for the computer to perform even in a limited way like human beings, detailed instruction sets with logical end-points at each node need to be created. It requires the imagination and creativity of the human mind to do this programming in a really effective manner - i.e. the computer actually being able to do what the human being in the same position would have been able to do. Also, it requires human ingenuity to engineer the machine to do this efficiently - within the desired speed and operating cost constraints. This role of an accelarator or an enabler of the "outsourcing" of hitherto human performed activity to machines will be increasingly in demand over the next 10-15 years.

This role will require a unique mix of skills. First and foremost, the role requires a detailed understanding of business processes, the roles played by the various players, the inputs and outputs at various stages. The business process understanding needs to span multiple companies and industries. Let's take something that Clark refers in his article: change to a flight reservation. The business process calls for not just access to the reservations database and the flights database, but also things like changing meal options, providing seating information (with information about the aircraft seating chart), reconfirming the frequent flyer account number, etc. Additionally, providing options for payment if there is going to be a fee involved.

Second, the role requires the ability to understand the capabilities of IT platforms and packages to able to perform the desired function. This role actually has two components. One is the mapping of human actions into the logic understood by a computer system. The second is the system architecture/ engineering side, which is the configuration of the various building blocks (comprised of different IT "boxes" delivering different functionality) to create an end-to-end process delivery capability. Given the lack of standards that exist for these types of solutions, any deep skills in this area involves understanding the peculiarities of specific solutions in a great level of detail.

I'd love to hear more from readers on this. Have you seen these roles emerging in your industry? What other types of skills does the enabler or accelarator role need?

Monday, August 3, 2009

Why individual level data analysis is difficult

I recently completed a piece of data analysis using individual level data. The project was one of the more challenging pieces of analysis I have undertaken at work, and I was (happily, for myself and everyone else who worked on it) able to turn it into something useful. And there were some good lessons learned at the end of it all, which I want to share in today's post.

So, what is unique and interesting about individual level data and why is it valuable?
- With any dataset that you want to derive insights from, there are a number of attributes about the unit being analyzed (individual, group, state, country) and one or more attributes that you are trying to explain. Let's call the first group predictors and the second group target variables. Individual level data has a much wider range of predictor and target variables. There is also a much wider range of interactions between these various predictors. For example, while on an average, older people tend to be wealthier, individual level data reveals that there are older people who are broke and younger people who are millionaires. As a result of these wide ranges of data and the different types of interactions between these variables (H-L, M-M, M-H, L-H ... you get the picture), it is possible to understand fundamental relationships between the predictors and the targets and interactions between the predictors. Digging a little deeper into the people vs wealth data, what this might tell you is that what really matters for your level of wealth is your education levels, the type of job you do, etc. This level of variation is not available with the group level data. In other words, the group level data is just not as rich.
- Now, along with the upside comes downside. The richness of the individual level predictors means that data occassionally is messy. What is messy? Messy means having wrong values at an individual level, sometimes missing or null values at an individual level. At a group level, many of the mistakes average themselves out, especially if the errors are distributed evenly around zero. But at the individual levels, the likelihood of errors has to be managed as part of the analysis work. With missing data, the challenge is magnified. Is missing data truly missing? Or did it get dropped during some data gathering step? Is there something systematic to missing data, or is it random? Should missing data be treated as missing or should it be imputed to some average value? Or should it be imputed to a most likely value? These are all things that can materially impact the analysis and therefore should be given due consideration.

Now to the analysis itself. What were some of my important lessons?
- Problem formulation has to be crystal clear and that in turn should drive the technique.
Problem formulation is the most important step of the problem solving exercise. What are we trying to do with the data? Are we trying to build a predictive model with all the data? Are we examining interactions between predictors? Are we studying the relationship between predictors one at a time and the target? All of these outcomes require different analytical approaches. Sometimes, analysts learn a technique and then look around for a nail to hit. But judgment is needed to make sure the appropriate technique is used. The analyst needs to have the desire to learn to use an technique that he/she is not aware of. By the same token, discipline to use a simpler technique where appropriate.

- Spending time to understand the data is a big investment that is completely worth it.
You cannot spend too much time understanding the data. Let me repeat that for effect. You cannot spend too much time understanding the data. And I have come to realize that far from being a drudge, understanding the data is one of the most fulfilling and value added pieces of any type of analysis. The most interesting part of understanding data (for me) is the sheer number of data points that are located so far away from the mean or median of the sample. So if you are looking at people with mortgages and the average mortgage amount is $150,000, the number of cases where the mortgage amount exceeds $1,000,000 lends a completely new perspective of the type of people in your sample.

- Explaining the results in a well-rounded manner is a critical close-out at the end.
The result of a statistical analysis is usually a set of predictors which have met the criteria for significance. Or it could be a simple two variable correlation that is above a certain threshold. But whatever be the results of the analysis, it is important to base the analysis result in real-life insights that can be understood by the audience. So, if the insight reveals that people with large mortgages have a higher propensity to pay off their loans, further clarification will be useful around the income level of these people, their education levels, the types of jobs they hold, etc. All these ancillary data points are ways of closing out the profile of the "thing" that has been revealed by the analysis.

- Speed is critical to get the results in front of the right people before they lose interest.
And finally, if you are going to spend a lot of everyone's (and your) precious time doing a lot of the above, the results need to be driven in extra-short time for people to keep their interest in what you are doing. In today's information-saturated world, it only takes the next headline in the WSJ for people to start talking about something else. So, you need to basically do the analysis in a smart manner, and also it needs to be super-fast. Delivered yesterday, as the cliche goes.

In hindsight, it gives me an appreciation of why data analysis or statistical analysis using individual level data is one of the more challenging analytical exercises. And why it is so difficult to get it right.

Tuesday, July 21, 2009

More data visualization - this time about books

Ever wonder where the proof was about reading ..ummm, erotica being bad for you. Here it is. Check this link out.

An interesting study was done that went somewhat like this.
- Get the ten most frequent "favorite books" at every college using the college's Network Statistics page on Facebook. Possibly these books represent the intellectual calibre of the college.
- Get their SAT/ACT scores for the colleges.
- You can now get a relationship between types of book read and scholastic achievement

The results are pretty impressive, though still somewhat dubious. According to the study, Classics is usually good for you (agree with that), Erotica is way bad. Controversially, so is African-American literature and chick-lit. In the link, check out the visual that stacks the book by genre.

Make what you want about this, but be careful between causality and correlation.

Saturday, July 18, 2009

Data visualization


An example of a really well-done graphic is from the NOAA website. Science and particularly math afficionados seem to have a particular affinity to following weather science. (I am wondering whether it is a visceral reaction to global warming naysayers who, the scientists think, are possibly insulting their learning.)

The graph is a world temperature graph and this type of graph has come in so many different forms, it is difficult not to have seen such a graph. What I like about this is the elegant and non-intrusive form in which the overlays are done.
• By using dots and varying the size of the dots, the creator of the graph is making sure that the underlying geographic details (important in a world map where there is great detail that needs to be captured in a small area, therefore you cannot use very thick lines for country borders) still come through.
• The other thing that I liked is some of the simplications the creator has made. The dots are equally spaced but I am pretty sure that’s exactly not how the data was gathered. But to tell the story, that detail is not as important.

The graphic came from Jeff Masters' weather blog which is one of the best of its kind. Here's a link if you are interested.

Wednesday, July 15, 2009

Two great finds for physics fans

Back after a long break in the posts. Call it a mixture of home responsibilities, writer's block and just some plain old laziness.

One of my other interests (apart from statistics and social science) is physics and technology. I really enjoy reading about emerging applications of technology in various spheres of social and economic importance. The Technology Quarterly of the Economist is one of my treasured reads (though I end up reading very little of it, because of me wanting to leave aside "quality time" to do the reading).

I want to share two recent finds in the science and physics space. One is a really good book called "The Great Equations" by Robert Crease. The book covers ten of the seminal equations in physics and basically spins a story around how the equation formulator came about to creating the equation. There is usually a little mathematical proof behind the story usually, but most of the book is about the professional journey made by the scientist from an existing view of the world (or an older paradigm, to be more exact) to a new paradigm. And the paradigm is usually encapsulated in the form of an equation.

I found a couple of aspects about the journey extremely interesting. One, it was fascinating to have a window into the minds of physics greats (Newton, Maxwell, Einstein, Schrodinger, to name a few) and see how they synthesized the various different world views around them to create or arrive at their respective equations. The ability to deal with all the complexity of observed phenomena, the different philosophies and world views and to come up with something as elegant as a great equation, that defines genius for me. The second aspect that I found extremely interesting was that there was usually years and years of experimentation or mathematical work that preceded arriving at the great equation. One might be inclined to think that the great equations (given their utter simplicity) happen through a flash of inspiration. Nothing could be further from the truth.

The next find were the Feynman lectures. Now, many of us have read some of the Feynman lectures or have seen the lectures on a place like Youtube. But how cool would it be to have these lectures be annotated by Bill Gates? Check this link out at the Microsoft Research website. And happy watching!

I am guessing this blog has a fair share of aspiring or one-time physics and engineering fans. How do you keep your engineering bone tickled? I'd love to hear your pet indulgences.

Wednesday, July 8, 2009

Market chills

I have argued in a number of recent posts: here, here and here that we are nowhere close to the bottom when it comes to this economic downturn. The jobless numbers are back to sliding downwards at an accelerated pace after one month of deceleration.

And the markets seem to have caught the chills.

We discussed this at work a few months back. Someone who is very well-respected in banking circles and who has seen a few past recessions called out that you can tell that a recovery is underway when there is a sustained period where the indicators yo-yo between good and bad news. We seem to be entering this phase now.

Friday, July 3, 2009

Best Coffee Survey and a research methodology question

A recent Zagat survey rated the best coffee in the US. The best coffee rating went (expectedly, I guess) to Starbucks. Even though I have had better coffee at other places, I guess Starbucks combines great coffee with ubiquitous presence and therefore ends up getting the top rating. Now, I think Starbucks coffee is good and the baristas are extremely friendly, but in terms of pure coffee flavour, I would rate Panera's Hazelnut coffee higher. Also some of the Kona coffees that you find at places like WaWa are also really good. Any kind of place serving Jamaica's Blue Mountain coffee will obviously be great. So what makes Starbucks special? Are there other factors at play beyond the pure taste of the coffee.

One hypothesis is that the national-level presence of Starbucks could be contributing towards the voting going for Starbucks. In places where Starbucks has to compete with other chains like Peets (San Francisco) and Dunkin Donuts (New England), comparative ratings between Starbucks and other chains shows a narrower gap. In places where Starbucks has not competition however, it is likely to get disproportionately good ratings.

Let us say you are one of the contributors in the survey and are in St.Louis, MO. The competition for Starbucks in St.Louis is likely to be (I guess) the burnt robusta coffee at the local restaurant. In such a market, Starbucks will enjoy a clear advantage, both for the quality of the coffee as well as the ambience. So, let's say, you had to rate Starbucks on a scale of 1-5. It is likely you would give Starbucks a 4-5 in a non-competitive market, such as St.Louis, in the absence of valid benchmarks or competition to compare against. In a competitive market dominated by multiple brands, the difference between Starbucks and other brands is likely to be narrower. Also, the assertion can also be made that a more discerning audience (having had the opportunity to sample multiple chains) is less likely to give extremely high scores (4s and 5s) to any of the choices under consideration.

Therefore, the sampling design and the analysis methodology becomes extremely critical for surveys around this. To avoid the "no-competition" bias, there could a number of questions a market research analyst would need to ask herself:
1. Should we use only data points from places where there are multiple chains in the same geography? (Doesn't sound fair. We will be throwing away data, which a lot of sensible people have explained is a cardinal sin. We should probably weight the information in some way).
2. Should we consider data for the analysis only where a person has provided ratings about multiple chains voluntarily or penalize when people have not rated a chain that could have been rated?
3. Or are there modeling solutions available to manage this conundrum? Topic of my next post!

Sitemeter