It is easy to get fooled by how the statistics interpret data. Sometimes, analysis of big data sets lead to conclusions that may not make sense. Also, the cause and effect do not work quite the same when the big data analysis shows a correlation. Just because there is a correlation does not mean that there is a cause and effect. Take the example of Kaggle… they ran a contest in 2012 on the quality of used cars and the characteristics of those cars. A used car dealer supplied the data to predict which cars were likely to have problems, their characteristics and what were the other cars that were not so likely to have problems. A correlation analysis showed that cars painted orange were far less prone to have defects – about half the rate of other cars. What has the car color got to do with problems? Color has no correlation and rightly so – this was just the chance event that was pulled out. But once, such a correlation between the car defects and color had been found out, the conclusions that can be drawn tends to get ridiculous. • Paint your car orange to have fewer defects. • Buy a orange car and your car will last longer, no matter how you treat it and forget about the oil change. • If you have an orange car, then you do not need to maintain the car. However, these conclusions get more complicated the more you use them. Even with the most complicated analysis, it is important to think about reason rather than believe everything that can be concluded.
Similar Posts
Integrative genomics viewer
It is not possible to use simple plotting or charting programs to view the data and draw meaningful conclusions. A dedicated viewer that enables large data sets to be displayed correctly is one of the best options. It is open source with copyrights by Broad and U. CaliforniaIf you are looking at genomics data and…
Extensible Open source -omics software
Understanding complex data takes effort. This graphic shows co-morbidites in COVID-19 that was accomplished by a piece of software called Cytoscape The number of open source software that is available is a big list. One of them is Cytoscape. It has ability for wonderful integration of network data from various sources that can be analyzed…
Simulating the whole brain one cell at a time.
Dr Dharmendra Modha’s group at IBM has been simulating the whole brain one cell at a time. They started with cat brain simulation and have almost reached the whole brain simulation except it runs about 1542 times slower. This is a simulation of nearly 1.6 billion virtual neurons and 9 trillion synapses. The power consumption…
Virtual Reality in Biology
Virtual Reality (VR) is moving forward so rapidly that it appears that we will be using that as a primary means for visualizing almost all parameters. The part that is stunning is some of the code that is being developed to enable VR. This uses the existing technology and now uses it in applications where…
Enable drug discovery
Drug discovery is hard.Amazing to see the databases that are available for public access that enable drug discovery. Broad institute publishes The Connectivity map (CMAP)which is a database of gene signatures of transcriptional response to perturbation of many cell lines. This is incredible amount of data that is available in the public domain to be…
NIH thousand genome project
How do you access 200 Terabytes of data? NIH is collecting data of many genomes that are being stored as computer code. These are then stored for researchers to access for their purpose. Being NIH, these large data sets are available for free. But, how do you process this amount of data. About the only…