Showing posts with label regression. Show all posts
Showing posts with label regression. Show all posts

Tuesday, September 6, 2016

Location is Time or Money

Using data from over 175,000 rentals from the real estate listing service StreetEasy, the site FiveThirtyEight asks the question:
How much would you be willing to pay to shave a minute off your commute? For New Yorkers, the answer appears to be around $56 per month. That’s how much more New Yorkers pay in rent, on average, for a one-bedroom apartment that’s a minute closer by subway to Manhattan’s main business districts.
They plot median monthly rental versus commuting time to the 42nd Street and Chambers Street subway stations. They fit, what appear to be non-parametric regressions, curves to four different groups of rentals: studios, and 1- ,2- , and 3-bedroom apartments. As expected, rental prices fall with a longer commute. From the 1-bedroom curve they estimate that, on average, a one minute shorter commute costs around $56 more a month.

Monday, February 1, 2016

Learning The Alphabet

Software engineer Erik Bernhardsson took a sample of 50,000 fonts, with characters as varied as shown in the compilation above, and looked for basic underlying structure with a neural network. A neural network is statistically a linear combination of nonlinear functions of linear combinations of input variables. Here, the input variables are digital images of each font character expressed as vectors. Iterative adjustment, termed learning, is applied to produce a linear combination of the inputs. An output estimate of the input character is computed from the other set linear coefficients. All coefficients are chosen to minimize a measure of lack of fit. Bernhardsson then looked at the mean and median of the resulting output characters.
Mean of all the output fonts.
Median of all the output fonts.
Note how readable the mean and median fonts are, when the individual input fonts are extremely varied, as shown above in the first image. He goes on to interpolate fonts, apply random perturbations, and even generate new fonts by sampling from a multivariate normal distribution of the font vectors. 

A mean of a collection of fonts we have seen before using a technique of simple visual averaging.


Monday, March 23, 2015

You Move Me

Here's an example an interactive linear regression demonstration which is like many programs that, these days, are included with nearly all basic statistics textbooks. This one works well and it's easy to move around outliers as well as placing data points close to the mean, in both cases watching what happens to both the slope and the intercept. It makes for many more fun explorations. Via Statistical Modeling, Causal Inference, and Social Science.

Monday, October 20, 2014

Are You Un-fashionably Late?

Here is a histogram of when people showed up to a party. From FiveThirtyEight, begun by Walt Hickey, but then crowdsourced from readers. Of 803 guest times submitted, the median guest arrived almost an hour after the party's start time. Four guests showed up over 3 hours late. How fashionable is that?
They also looked at a scatterplot of arrival time against number of party guests.  They fit a regression, with only a 5%  R2, and the following interpretation: as host you should expect the mean guest to arrive 42 minutes after the party's start plus 4 minutes for every additional 10 guests. So comparing parties that differed by 10 guests, on average, the guests to the larger party arrived 4 minutes later.

Monday, April 29, 2013

Kernel Pinterest

Here is a nice idea for displaying a bit more than summary statistics on the variables included in regression studies. This is from the paper "I Need to Try This!": A Statistical Overview of Pinterest. Pinterest is a pin-board photo sharing website. Among other things, this study models the number of re-pins of a given photo with a Negative-Binomial regression.

The table above shows the medians, means, and maxima for non-negative count data included in the regressions. The minima are all zero. Along with these summary statistics are small thumb-nail kernel density estimates of the distributions of the variables. Now granted the variables involved take on only integer values and these distribution curves are continuous, but it is much better than the usual limited summary statistics, shown below, that are often given in other regression studies.


Monday, April 1, 2013

Salty Residuals

Washington, this Winter, has not been wearing "on his smiling face a dream of Spring". On the contrary, Spring has been continually cold and damp much like the movie Groundhog Day. Just last week we had an unusual mid-March snowstorm. Groundskeepers spread salt along campus walkways to speed the snow melt. Their spreader sprayed the salt left and right as they drove down the path. After melt, the salt remained attracting moisture, absorbing - not reflecting - the light that fell on it, leaving a dark, wet residual on the asphalt.

At every point down the walk we can see the horizontal spread of the dark residuals, much concentrated near the center of the walk with lesser concentrations to the right and left. Horizontally, what remains is a bell-shaped distribution of salt deposition. The distribution has this same, consistent shape at each place down the walk. In this angled, perspective view the distributions look more skewed to the right. But the symmetry can be seen more accurately in the rotated image below.
This rotated image now shows the distributions in vertical slices as we move horizontally along the walk. The distributions are centered along the same horizontal line with the same shape and degree of vertical spread.

This is exactly the image of ideal residuals from a simple linear regression fit to data plotted against an explanatory variable from a uniform design. Of course, a different design placement of the explanatory variable would vary the pattern horizontally, but not so vertically. Other design patterns could arise from the spreader moving faster or slower down path leaving a more uneven, erratic deposition of salt - more at some steps along the walk than at others. But the assumptions for such a regression model still require identical, vertical normal distributions of scatter around a straight line of means irrespective of the horizontal position down the path. For an even, uniform walk down the path, our salty residuals model and reflect the ideal behavior of regression residuals.

Sunday, March 13, 2011

Multiple Regression Model c. 1911



A multiple regression model from George Undy Yule's "An Introduction to the Theory of Statistics" (1911) page 242. The residuals can be seen in the edge-on view.