Monday, March 17, 2014

In Honor of Pi Day last Friday

In honor of Pi day this past Friday, March 14. I posted this long ago in 2008. It is still my favorite pie chart.

Monday, March 10, 2014

Conditional Probability

Here is a clever, interactive, simulation display of conditional probability from Victor Powell, via flowingdata.

Balls fall from the sky, uniformly distributed across the display window. Some hit a red shelf of adjustable width. Here it is set at a width so that
P(A) = 20% of all the falling balls hit the red shelf.
Another lower blue shelf, overlapping a bit with the red one, is set so that
P(B) = 12% of the balls hit the blue shelf.
A mere P(A and B) = 6% of the balls hit both shelves, indicated by the mixture of red and blue to get purple. But, as the simulation says,
"If we have a ball and we know it hit the red shelf, there's a 30.0% chance it also hit the blue shelf" and 
"If we have a ball and we know it hit the blue shelf, there's a 50.0% chance it also hit the red shelf". 

Below are connected bars showing, by their length, the color composition of the dropped balls. We can easily see how these proportions are obtained by visually estimating that fraction that purple makes up of the balls that hit the red shelf. That is the ratio, purple / (red + purple) = 30% or what fraction purple makes up of the balls that hit the blue shelf. This is the ratio, purple / (purple + blue) = 50%.

Monday, March 3, 2014

Dancing Statistics

A still image from the project Communicating Psychology to the Public through Dance, produced by Lucy Irving, Elise Phillips, and Andy Field supported by the British Psychological Association and IdeasTap. Four videos: my favorite Frequency Distributions, Sampling and Standard Error, Variance, and Correlation.

In this image from the first of the videos the dancers start our in one large unorganized group, some dancing with very slow movements, some with very quick movements, and as one would then expect more with movements of a more intermediate speed. As they dance they sort themselves out, from the slower movement dancers on the left, to the more rapidly moving dancers on the right, building up a sample from a bell-shaped distribution. Very clever.

There is a video about correlation with dancers performing the same movements together or nearly opposite movements together. As they mention in the text of the video, these movements are just co-occurrences, one does not cause the other: correlation is not causation.

There is a video about variation with dancers performing variations on the same set of movements.

Another is about sampling and standard error. In this dance, a single blue-shirted dancer performs his movements to indicate the four corners of a rectangle. He and the rectangle defined by his  movements are termed the population. Then several red-shirted dancers mark four corners in their own styles producing various quadrilaterals that estimate the rectangular shape of the blue-shirted dancer.

I do think calling that initial, single blue-shirted dancer a population could be misleading, especially since at the beginning of this video the text mentions "a large group (a population)". This, of course, is the usual view: the large group is the population from which we observe samples to estimate it. But perhaps better, in this setting, would be to talk more generally about a statistical model. This is a model for a dancer's movements. These movements depend on the physical aspects of the dancer: height, limb length, reach, flexibility, etc. They also depend on artistic intent, style, technique, etc.

The blue-shirted dancer specifies the results of a certain collection of all of these aspects. This becomes a parameter, a target. The red-shirted dancers sample from the model of movements to estimate this parameter. The variety and range of their movements display the sampling variability as they attempt to match the governing shape (parameter) of the blue-shirted dancer. This view is more general than the viewing of sampling as from a fixed large group population.

Monday, February 24, 2014

Pop to Area

Map of the US with population switched for physical size. This graphically shows the result of a sorting: reassigning the largest state name by population (California) to the largest state by area (Alaska) and so on down to assigning the smallest state name by population (Wyoming) to the smallest state by area (Rhode Island). Of course, part of the effect is lost since Alaska is much larger than portrayed: it is more than 1/5 of the area of the contiguous US. Still this is a fun graphic.



Monday, February 17, 2014

Most Pleasant Places

Designer and software engineer Kelly Norton has produced a map to find the greatest number of pleasant days in the year around the continental US. Her definition:

“pleasant” here means the mean temperature was between (55° F and 75° F), the minimum temperature was above 45° F, the maximum temperature was below 85° F and there was no significant precipitation or snow depth.
You can roll the cursor over the map and reveal bar charts of the number of pleasant days by month. Although it may be positively correlated with temperature, I think a big omission from this definition of "pleasant" is a measure of humidity.

kellegous.com via flowingdata.com

Monday, February 10, 2014

"Statistical Analysis: It's So Beautiful"

Words that should be heard daily (a Moneyball reference). 

Monday, February 3, 2014

A Real Stem and Leaf Display

Here is a real stem-and-leaf display from a collection of deconstructions and re-organizations in the book Art of Clean Up by artist/designer Ursus Wehrli. Many more here.