It is easy to get fooled by how the statistics interpret data. Sometimes, analysis of big data sets lead to conclusions that may not make sense. Also, the cause and effect do not work quite the same when the big data analysis shows a correlation. Just because there is a correlation does not mean that there is a cause and effect. Take the example of Kaggle… they ran a contest in 2012 on the quality of used cars and the characteristics of those cars. A used car dealer supplied the data to predict which cars were likely to have problems, their characteristics and what were the other cars that were not so likely to have problems. A correlation analysis showed that cars painted orange were far less prone to have defects – about half the rate of other cars. What has the car color got to do with problems? Color has no correlation and rightly so – this was just the chance event that was pulled out. But once, such a correlation between the car defects and color had been found out, the conclusions that can be drawn tends to get ridiculous. • Paint your car orange to have fewer defects. • Buy a orange car and your car will last longer, no matter how you treat it and forget about the oil change. • If you have an orange car, then you do not need to maintain the car. However, these conclusions get more complicated the more you use them. Even with the most complicated analysis, it is important to think about reason rather than believe everything that can be concluded.
Similar Posts
PySB: Biological models in Python
As scientists make biological models, it is difficult to create frameworks of mathematics and formula’s around it. Usually this requires significant customization for each model. Dr Sorger’ Lab has developed a Python framework for creating systems biology models. These translate the equations to the models and allow a tedious task to be significantly simplified by…
Pollen robots
In spring, the flowers and trees blooming look very attractive but also signal a release of pollen which triggers allergies and hay fever in many people. Many drugs have become available to combat the symptoms but the release of pollen is significant. There is not much that can be done to control the amount of…
Know the immeasurable in data mining
Data mining has several methods that analyze the data and come to conclusions on the data. The tools are a mixture of statistics, mathematics and reasoning. The tools are very useful and they help in come to conclusions very well. So if you are in the business of selling a commodity to the youth through…
Cronbach’s alpha
A statistics concept that has been used for a specific purpose. This has been used to interpret how well the scales, surveys (or test items) function for determining a factor under consideration. It is used to determine reliability or the internal consistency of the test item. The surveys are used to measure things that are…
Radiopharmaceuticals
Development is complex: Usually are highly cancer specific, high persistent expression throughout the tumor and metastasis, readily accessible on the cell surface, not shed and thoroughly validated. Many isotopes are possible and many targets possible such as Nectin4, DLL3, FAP, Somatostatin receptors, PSMA, Integrins, CXCR4, Claudin 18.2, Norepinephrine transporter, EGFR2. Clearance from tumor will matter…
Walmart : Data mining in retail
Walmart as is well known is a big retailer, not just big but massive. They have been a big retailer since a long time and have been processing large data before the term “Big data existed”. The interesting feature about them is that they are about the biggest retailer with a net turnover of $450…
