Showing posts with label variance. Show all posts
Showing posts with label variance. Show all posts

Monday, November 3, 2014

The Curse

The Curse of Dimensionality addresses the difficulty of dealing with multivariate data. It warns us that, for a set of data in high dimensions, local neighborhoods are almost certainly empty of data points and neighborhoods that are not empty are almost certainly not local. 

In explaining this result, biostatistician Jeff Leek thought of clever a demonstration and got his graduate student Prasad Patil to build an interactive program to illustrate the Curse. In the screen shots above, samples of 100 points are randomly and uniformly generated in 1,2,3, and 4 dimensions in the unit cube. Subsets are examined in cubes with edge length of 1/2. In 1-dimension, the simulation contained 55% of the data in a line segment of length 1/2 (expected is, of course, 50%). In 2-dimensions, the simulation contained 31 % in a square with sides of length 1/2 (expected is 25%). In 3-dimensions, the simulation contained 14% in a cube with sides of length 1/2 (expected is 12.5%). And finally, the simulation contained just 4% of the data in a 4-D cube with sides of length 1/2 (expected is 6.25%). As the dimension grows, smaller and smaller percentages of the data can be found in regions with linear dimensions, that our low dimensional intuition tells us, are not small. Balancing the variance of a large neighborhood with the low bias of a small neighborhood is incredibly difficult in high dimensions.

This is a nice way to help visualize the Curse.

Monday, July 21, 2014

Adding Economic Noise

Two months ago the New York Times had a very informative visualization of monthly economic data with added variability due to sampling error. In the above screen shot we can see how repeated sampling variability can change a steady job growth graph to many different shapes of what it might have looked like with repeated sampling.
Here is another, an accelerating job growth graph and one possible result of the same graph with sampling variability showing very stable job growth, via Statistical Modeling ... One commenter there notes how easy it is for us to reading meaning into random noise.





Monday, April 28, 2014

Conditioned Steps

Conditional frequency distributions of footfall wear on several wooden steps at the Simon Pearce glassworks Mill at Quechee, Vermont. The lowest of these four steps is on the left, showing a distribution of wear ranging across much of the step. The variability of footfall placement along the edge of this narrow step is great. The next step to the right in this image shows a greater concentration of wear near the middle of a slightly wider step, allowing for a more full foot placement. The next step is wider still with a concentration of wear shifting slightly to the right (downward in this image). Finally, arriving on the top step of the landing (rightmost in this image) the footfalls cause a  nearly circular pattern of wear as feet need to turn to the right (where you can see my feet standing to take the picture). This path continues ascending on the next set of stairs to an upper floor.

Left to right we see conditional frequency distributions of wear: conditioned on each step. Imagining a line connecting the means of these distributions would show a line of decreasing slope. This line is that of the conditional mean of footfall placement conditioned on each step. The changing variability of wear on each step, shows the pattern of the conditional variance: greater variance on the lower (leftmost) step and lesser variance on the higher (rightmost) step, a concept termed heteroscedasticity.