Showing posts with label causation. Show all posts
Showing posts with label causation. Show all posts

Monday, January 4, 2016

New Year's Resolution: Spot Some Bad Science

 
Here's your assignment: using one (or more) of the 12 methods in the infographic above spot some bad science this year (from COMPoUND iNTEREST download their graphic here). Let me know your findings. Enjoy and Happy New Year. Via datavizblog.

Monday, February 9, 2015

Selective Correlation

Sociologist Gabriel Rossman of UCLA has an article "When Correlation Is Not Causation, But Something Much More Screwy" in The Atlantic. It graphically shows how calculating correlations on selected samples can produce misleading results. He imagines two characteristics of aspiring actors, ability (mind) and attractiveness (body) plotted in a scatterplot. Random observations from independent standard normal distributions are drawn to represent these variables in a population of actors. Most actors are centered around zero with fewer well above or well below, on mind or body. But since mind and body are chosen independently there is no correlation shown in the plot, that is, no tendency of the plot to tilt with a positive or negative slope. 

But then he imagines computing the correlation between mind and body from a sample of working actors. To get work, his actors have been selected by casting directors only if they have a high value for the sum of both mind and body. He has marked these working, observed actors with small triangles in the plot. The remaining aspiring non-working and unobserved actors are marked with small circles. The plotted pattern of triangles for the working, observed actors has a definite negative correlation suggesting, wrongly, that either the more able actors are not much to look at or the most attractive actors are dunces in their acting ability. Although this might fit some stereotypical caricatures, it is entirely due to the selective sampling based on the definition that produced our observed sample. He continues with an SAT example as well.

His scatterplot can be improved. Measurements with equal variability should be plotted in a square plot. The plotting characters for the unobserved and observed actors could be more pronounced. See an example below.


Monday, November 17, 2014

Data Literacy: It's Elementary


The Washington Post has an article this morning (17 November, 2014): "In elementary schools, lessons on data literacy," by IT reporter Mohana Ravindranath. She describes a "growing movement of educators creating lesson plans to teach students to collect and analyze data." One goal is "to derive opinions from measurable, real-world data." Another, is to address the shortage of "managers and analysts who can make decisions based on big data analysis," according to management researcher Michael Chui. The Washington Post article goes on to quote Chui:
“It makes sense for us to be thinking about education, starting in early childhood, about concepts such as the difference between correlation and causation, what it means to have a bias as you think about data, conditional probability. These are things we as humans don’t naturally do . . . these are learned [concepts],” Chui said in an interview. He added that curricula should teach students about the realistic limitations of data sets — extraneous information, or sampling error, for instance.
The article describes students collecting their own data. Third-grade students collect daily temperature data, fifth-grade students record the hours of daylight and relate them to the earth's motions, and even kindergarten children "recording predictions for whether it will be sunny outside the next day, or which foods will decompose fastest, along with the results."

Says one science coordinator at an elementary school, evaluating the effectiveness of these lessons is "ultimately if the kid’s able to have a conversation about it and ask questions about it.”

A great goal for students of all ages. That this is taught and expected of even elementary school students is inspiring.

(On a very minor display note: the introductory graphic to this story is an image of a computer monitor showing results from a school's Science Festival using software from Tuva Labs. Dot plots are displayed showing the arm spans by gender. I wonder about the zoom-in that is shown for one data point. It seems only to extract the same dot plot that's on the screen. That's something to ask a question about!)

Monday, May 19, 2014

Nonsense! Correlations!

In 1926, G. Udny Yule published a paper in the Journal of the Royal Statistical Society, titled, "Why do we Sometimes get Nonsense-Correlations between Time-Series?--A Study in Sampling and the Nature of Time-Series" He begins by describing the problem:
The graph above illustrates this problem for the time series of  mortality (death rate) and marriages (marriage rate) in England from 1866 to 1911. Since these two time series are both decreasing over these years, they start out high together and end up low together, so their correlation coefficient is high (0.95). Their scatterplot is shown below, and if these data were from independent observations, such a high correlation could be quite informative. But since these series are correlated in time, this correlation likely only reflects their time trends and not any other connection. Such a correlation is called spurious.
More recently, a friend (thanks JCT) pointed me to work of Tyler Vigen, explained very well in his video. He has written a program to find such nonsense correlations with hilarious results, such as the high correlation between the per capita consumption of cheese in the US and the number of people who died by becoming entangled in their bedsheets.
This is great fun, and his work has been widely circulated on the internet lately, here, here, and here as examples. Many of these discussions emphasize that "correlation does not imply causation" and some even invoke this xkcd cartoon

But spurious correlations can be more insidious than these graphs illustrate. These time series only illustrate how correlating time related variables can be misleading, but spurious correlation can arise with data not correlated in time. For example, drunk driving and fatal accidents increased more in localities that had banned smoking in bars and restaurants than those that had not (from this paper). One might be tempted to conclude that banning smoking causes these bad outcomes. But something more is at work here. Smokers needed to travel farther to find alternative localities that had not banned their addiction. They were on the road more and therefore had more accidents. In economics, these often referred to as an unintended consequences. Many other settings have similar lurking or confounding variables or behaviors, that have nothing to do with time series, that can also give rise to spurious correlations.

But if we have a correlation for two time series, how might we determine whether it is spurious or not? One approach for the Yule's marriage and mortality data is the following: If there is a real, causal connection between mortality and marriages, then when mortality changes (year-to-year), we should see a related change in marriages (year-to-year). If these changes are not related, then the correlation between mortality and marriages (0.95) is likely spurious and misleading. Below is a scatterplot of the year-to-year differences of mortality and marriages. Their correlation is 0.064, suggesting that the previous, time-dependent, correlation of 0.95 is spurious and misleading.