It is easy to get fooled by how the statistics interpret data. Sometimes, analysis of big data sets lead to conclusions that may not make sense. Also, the cause and effect do not work quite the same when the big data analysis shows a correlation. Just because there is a correlation does not mean that there is a cause and effect. Take the example of Kaggle… they ran a contest in 2012 on the quality of used cars and the characteristics of those cars. A used car dealer supplied the data to predict which cars were likely to have problems, their characteristics and what were the other cars that were not so likely to have problems. A correlation analysis showed that cars painted orange were far less prone to have defects – about half the rate of other cars. What has the car color got to do with problems? Color has no correlation and rightly so – this was just the chance event that was pulled out. But once, such a correlation between the car defects and color had been found out, the conclusions that can be drawn tends to get ridiculous. • Paint your car orange to have fewer defects. • Buy a orange car and your car will last longer, no matter how you treat it and forget about the oil change. • If you have an orange car, then you do not need to maintain the car. However, these conclusions get more complicated the more you use them. Even with the most complicated analysis, it is important to think about reason rather than believe everything that can be concluded.
Similar Posts
McGurk effect – you actually hear only what you see or want to hear!
The video below explains it all. But essentially our perception is dependent on what we see. Wonder how many of these effects exist without us realizing it {youtube}FefFfvriAwQ{/youtube}
Examples on how data mining helps with real world problems – data scientist
Facebook’s Jeff Hammerbacher is an interesting person. He literally coined the term – data scientist and has been one of the big proponents of data mining. He found that the big predictor whether people will take action is dependent on whether they had seen their friends do the same thing. This is true whether you…
Personalized medicine for treatment of Cancer.
In treatment of cancer, there has been a significant thought in the value of personalized medicine. In general there is agreement that each cancer is slightly different and therapies need to be targeted specific for different patients. However, since each combination of therapy is unique, they have to be customized for the patients. This works…
Mathematics prizes
There have been a variety of prizes to solve problems that are practical such as space rocket. Some companies have made a business out of problem solving through mediating solvers and problem generators – such as those mediated by Innocentive. However, rarely are problems solved just for the sake of solution of the problem. One…
Simulating the whole brain one cell at a time.
Dr Dharmendra Modha’s group at IBM has been simulating the whole brain one cell at a time. They started with cat brain simulation and have almost reached the whole brain simulation except it runs about 1542 times slower. This is a simulation of nearly 1.6 billion virtual neurons and 9 trillion synapses. The power consumption…
Labor statistics the Big Data way
Typically, labor statistics are collected tediously by the staff of Bureau of Labor Statistics by visiting, faxing and calling different stores, offices and online retailers in 90 cities across the nation and getting back nearly 80,000 prices for different items. This costs about $250MM per year and takes at least several weeks to put together….