Showing posts with label Poisson. Show all posts
Showing posts with label Poisson. Show all posts

Monday, June 3, 2013

Testing Poissonness Petals


Last week’s image was of cherry blossom petals that had fallen on a stone walk, a random realization of a spatial Poisson process. In such a process, the probability of some number of petals falling on any stone is proportional to the area of the stone. The mean number of petals falling on the larger stones is 8.71. The mean number of petals falling on the smaller stones is 5.58. Their ratio is 1.56. The ratio of the stones’ area is 1.5. Such a close agreement is what a Poisson process should produce.

Not so fast, implies commenter Kevin. He suggests that we should test whether our observed ratio of 1.56 is significantly different from the ratio of the stones’ area of 1.5. My intuition tells me that, with the large variability in a Poisson process and such a small sample of fallen petals, we would need a much larger difference between the observed and expected ratios for the difference to be significant. But let’s test it.

First, consider the larger stones. If we count those laid down horizontally, left to right, from top to bottom, my counts (again likely prone to error) are 12,9,10,9,11,13,3,13,11,7. For the larger stones laid down vertically, left to right, from top to bottom, I count 12,6,7,6,11,6,6,10,10,4,7 petals. For the smaller, square stones, left to right, from top to bottom, I count 3,3,2,6,8,4,8,11,4,5,7,6 petals.  Let’s investigate whether these data can be modeled with a Poisson distribution. For example, the mean and variance for the larger stones are 8.71 and 8.61, respectively, values close to the equal mean and variance expected for a Poisson distribution. But we can do more.

The image above is called a Poissonness plot (Hoaglin 1980). This one is for the sample of petals from the larger stones. It allows us to graphically test whether a Poisson distribution is an appropriate model for the petal counts. Let x[k] be the count of stones collecting k petals, we plot on the vertical axis: log(x[k])+log(k!) against k (on the horizontal axis). In such a plot, Poisson counts fall upon a straight line with a slope equal to the logarithm of the Poisson mean. This is what we see for the larger stones. The plot for the smaller stones is also straight.

Kevin suggests testing our observed ratio to see if it is significantly different from the expected ratio of 1.5. Since we are dealing with discrete distributions we can compute the exact probability distribution for this ratio. The larger stones have a length of 9.375. The shorter, square stones have a length of 6.25, for a ratio of 1.5. We take these as the null parameters of two independent Poisson distributions. Our samples of sizes 21 and 12 have sums that are also Poisson. The joint distribution is simply their product, allowing easy computation of the probabilities for the ratio of the means. Its discrete probability distribution is shown below (ignoring the small probability that the denominator is zero, 2.6 x 10-33 ).

Although this appears to be a darkly colored continuous probability density, it is actually a discrete distribution of probability mass plotted as densely packed individual vertical lines or spikes of probability. For this null distribution our observed ratio of means, 1.56, has an upper tail p-value of 0.39.  This, of course, is not significantly different from 1.5. In fact, our observed ratio would have to exceed 1.89 to be significant at the 5% level. Thanks for the prompting, Kevin.

Monday, May 27, 2013

Poisson Petals

Today is Memorial Day, a day to remember those that have fallen during US military service. It is also the traditional beginning of Summer. But Spring has been hard pressed to give-way to Summer here in Washington. Just two days ago. the temperature struggled upward, but stayed in the 60's (F), [10-15 Celsius]. A cold front kept it cold and rainy all day, much like much of our Spring this year.

Washington's iconic sign of Spring, the cherry blossoms, have long faded and fallen. This picture of fallen blossoms was taken a few weeks ago beneath a cherry tree that stretches like an umbrella above my front walk. What the picture shows is a realization of a spatial Poisson process. Such a random process counts, in continuous time, the number of petals that fall into non-overlapping regions. As the petals randomly land, the number of petals landing in any two separate paving stones are independent of each other. This would indicate that one petal, or its method or path falling from the tree, does not affect any other. The probability distribution of the count of petals on any paving stone depends only on the area of the stone.

I counted (likely with some error) the number of petals on each whole paving stone shown in this image. The mean number of petals on the square stones is 5.58. The mean number of petals on the rectangular stones is 8.71. If these were a result of a Poisson process these means should be proportional to the areas of the stones. The rectangular, larger stones are half again larger than the square ones, that is, the ratio of the areas (rectangular/square) is 1.5. Under a Poisson process we should expect the same for the mean. And sure enough the ratio of the means is 8.71/5.58 = 1.56.

Monday, August 8, 2011

Our Video

Award for Best Evidence of Inspiring Students at the 2011 Joint Statistical Meetings in Miami Beach. Double click on video for full screen.

Thursday, February 24, 2011

"Mnemonic" is itself a (Poisson) mnemonic


See The Futility Closet. I can now use the Poisson distribution to help me remember how to spell "mnemonic"!

Tuesday, July 3, 2007

Cycle Man


On Friday, June 29, 2007 Aubrey Huff first baseman for the Baltimore Orioles hit for the cycle. In one game, he had a single, a double, a triple, and a homerun. This has happened 276 times in major league history, only three times for the Baltimore Orioles and it is the first time it has happened at home in Baltimore and at Oriole Park at Camden Yards. In the major leagues he is the third player to hit for the cycle in 2007. So what is the distribution of such an event?

Interestingly, just this year Huber and Glen have modeled the distribution of such rare events in their article, “Modeling Rare Baseball Events – Are They Memoryless?” in the Journal of Statistics Education. They model three baseball events as a Poisson Process: no-hit games, triple plays, and hitting for the cycle. Using data from 1901 through 2004 they get the distribution shown above of how many cycles are seen in a given season. The observed results are a bit short on seasons with exactly 2 players hitting for the cycle, and there are a few more seasons than expected with exactly 1 player hitting for the cycle. The expected results are shown in red: a Poisson distribution with the observed mean of 2.19.

Over the time period examined, nearly 160,000 games were played, and only 0.14% of these games had a player hit for the cycle. As Huber and Glen found the cycle is rarer than a triple play but not as rare as a no-hitter.

So how closely does this match a Poisson process? One measure is a Poissonness Plot developed by Hoaglin (American Statistician 1980). If f[k] represents the observed frequency of a random count k then the log of the Poisson density would essentially be -λ+klog(λ)-log(k!). So that in a plot of log(f[k])+log(k!) against k, a straight line would indicate a good fit to a Poisson distribution. The slope of this line is an estimate of log(λ). Here log(λ) is 0.9583, so that an estimate of λ is 2.607. Of course the maximum likelihood estimate is simply the mean number of cycles per season 2.19. Since the variance of a Poisson distribution is also equal to λ, this would provide another estimate of 2.90. Note the Poissonness plot estimate falls in between the two.

Another way of checking for a Poisson process is to check the inter-arrival times, that is, the number of games between any two players hitting for the cycle. These inter-arrival times should follow an exponential distribution. The figure below shows the cumulative relative frequency of the observed time between cycles for the period 1901 through 2004. Also shown is the theoretical exponential cumulative probability distribution with a mean equal to the observed mean of 719.533 games. This indicates that the cycle process is memoryless. Even if there have been 1000 games without any major league player hitting for a cycle, you would still have to wait on average over 719 games more before one does.

For Aubrey Huff it was certainly a special night. Not only did he hit for the cycle he also had the 1000th major league hit and the 200th double of his career. Even with all that the Orioles still lost the game to the LA Angels 9 to 7.

Monday, June 11, 2007

Poisson Parking




The picture shows an aerial photograph, courtesy of the United States Geological Society, of the main visitor’s parking facility of business in Maryland on a Sunday morning in April. Look at the line of 13 parking spaces at the bottom of the picture. Notice the pattern in the oil stains leaked from the cars that park in those spaces. More oil is leaked in the spaces closest to the building.

Parking lots permit both the dynamic and static viewing of statistical processes. A time lapse film of customers entering and leaving a parking lot could allow us to estimate arrival rates, lengths of stay, or number of parked customers as time progresses. But automobiles, not being the cleanest of vehicles, leave their mark. This process is often modeled by a Poisson distribution. The actions of this process can be seen in static pictures as well.

Being Sunday no one is visiting the company, but the many previous visitors have left their marks. Notice the oil stains in the parking places. Much more oil is leaked and deposited in the places used most often. These places are the ones closest to the building. The one parking space closest to the building is a place for drivers with handicaps. Skip this space and examine the pattern starting with this first non-handicapped parking space. This first space will have oil stains when there are one or more cars in the lot. The second space will have oils stains when there are two or more cars in the lot, and so on. Thus, as the distance from the office front door to the car increases, then the amount of leaked oil decreases. This shows the steady state distribution of a multi-server queue. It takes the form of a truncated Poisson distribution, distributing Poisson probability only among the first few positive integers corresponding to the number of parking spaces.