Showing posts with label nonparametric. Show all posts
Showing posts with label nonparametric. Show all posts

Tuesday, September 6, 2016

Location is Time or Money

Using data from over 175,000 rentals from the real estate listing service StreetEasy, the site FiveThirtyEight asks the question:
How much would you be willing to pay to shave a minute off your commute? For New Yorkers, the answer appears to be around $56 per month. That’s how much more New Yorkers pay in rent, on average, for a one-bedroom apartment that’s a minute closer by subway to Manhattan’s main business districts.
They plot median monthly rental versus commuting time to the 42nd Street and Chambers Street subway stations. They fit, what appear to be non-parametric regressions, curves to four different groups of rentals: studios, and 1- ,2- , and 3-bedroom apartments. As expected, rental prices fall with a longer commute. From the 1-bedroom curve they estimate that, on average, a one minute shorter commute costs around $56 more a month.

Monday, February 1, 2016

Learning The Alphabet

Software engineer Erik Bernhardsson took a sample of 50,000 fonts, with characters as varied as shown in the compilation above, and looked for basic underlying structure with a neural network. A neural network is statistically a linear combination of nonlinear functions of linear combinations of input variables. Here, the input variables are digital images of each font character expressed as vectors. Iterative adjustment, termed learning, is applied to produce a linear combination of the inputs. An output estimate of the input character is computed from the other set linear coefficients. All coefficients are chosen to minimize a measure of lack of fit. Bernhardsson then looked at the mean and median of the resulting output characters.
Mean of all the output fonts.
Median of all the output fonts.
Note how readable the mean and median fonts are, when the individual input fonts are extremely varied, as shown above in the first image. He goes on to interpolate fonts, apply random perturbations, and even generate new fonts by sampling from a multivariate normal distribution of the font vectors. 

A mean of a collection of fonts we have seen before using a technique of simple visual averaging.