Showing posts with label sampling distribution. Show all posts
Showing posts with label sampling distribution. Show all posts

Monday, November 17, 2014

Data Literacy: It's Elementary


The Washington Post has an article this morning (17 November, 2014): "In elementary schools, lessons on data literacy," by IT reporter Mohana Ravindranath. She describes a "growing movement of educators creating lesson plans to teach students to collect and analyze data." One goal is "to derive opinions from measurable, real-world data." Another, is to address the shortage of "managers and analysts who can make decisions based on big data analysis," according to management researcher Michael Chui. The Washington Post article goes on to quote Chui:
“It makes sense for us to be thinking about education, starting in early childhood, about concepts such as the difference between correlation and causation, what it means to have a bias as you think about data, conditional probability. These are things we as humans don’t naturally do . . . these are learned [concepts],” Chui said in an interview. He added that curricula should teach students about the realistic limitations of data sets — extraneous information, or sampling error, for instance.
The article describes students collecting their own data. Third-grade students collect daily temperature data, fifth-grade students record the hours of daylight and relate them to the earth's motions, and even kindergarten children "recording predictions for whether it will be sunny outside the next day, or which foods will decompose fastest, along with the results."

Says one science coordinator at an elementary school, evaluating the effectiveness of these lessons is "ultimately if the kid’s able to have a conversation about it and ask questions about it.”

A great goal for students of all ages. That this is taught and expected of even elementary school students is inspiring.

(On a very minor display note: the introductory graphic to this story is an image of a computer monitor showing results from a school's Science Festival using software from Tuva Labs. Dot plots are displayed showing the arm spans by gender. I wonder about the zoom-in that is shown for one data point. It seems only to extract the same dot plot that's on the screen. That's something to ask a question about!)

Monday, March 3, 2014

Dancing Statistics

A still image from the project Communicating Psychology to the Public through Dance, produced by Lucy Irving, Elise Phillips, and Andy Field supported by the British Psychological Association and IdeasTap. Four videos: my favorite Frequency Distributions, Sampling and Standard Error, Variance, and Correlation.

In this image from the first of the videos the dancers start our in one large unorganized group, some dancing with very slow movements, some with very quick movements, and as one would then expect more with movements of a more intermediate speed. As they dance they sort themselves out, from the slower movement dancers on the left, to the more rapidly moving dancers on the right, building up a sample from a bell-shaped distribution. Very clever.

There is a video about correlation with dancers performing the same movements together or nearly opposite movements together. As they mention in the text of the video, these movements are just co-occurrences, one does not cause the other: correlation is not causation.

There is a video about variation with dancers performing variations on the same set of movements.

Another is about sampling and standard error. In this dance, a single blue-shirted dancer performs his movements to indicate the four corners of a rectangle. He and the rectangle defined by his  movements are termed the population. Then several red-shirted dancers mark four corners in their own styles producing various quadrilaterals that estimate the rectangular shape of the blue-shirted dancer.

I do think calling that initial, single blue-shirted dancer a population could be misleading, especially since at the beginning of this video the text mentions "a large group (a population)". This, of course, is the usual view: the large group is the population from which we observe samples to estimate it. But perhaps better, in this setting, would be to talk more generally about a statistical model. This is a model for a dancer's movements. These movements depend on the physical aspects of the dancer: height, limb length, reach, flexibility, etc. They also depend on artistic intent, style, technique, etc.

The blue-shirted dancer specifies the results of a certain collection of all of these aspects. This becomes a parameter, a target. The red-shirted dancers sample from the model of movements to estimate this parameter. The variety and range of their movements display the sampling variability as they attempt to match the governing shape (parameter) of the blue-shirted dancer. This view is more general than the viewing of sampling as from a fixed large group population.

Monday, March 25, 2013

Age of the Oldest Person You Know


From a Prudential TV commercial that has people place large, round stickers on a number line to represent the age of the oldest person they know. This forms a histogram or dotplot. The message is to have Prudential prepare you to have adequate money for all these years. You can add your own sticker at the Prudential website. Here is a video behind the scenes.