A lopsided population
Ask the students of a school how many library books they borrowed last month. Suppose 40% borrowed none, 30% borrowed 1, 15% borrowed 2, 10% borrowed 3 and 5% borrowed 4.
This distribution is nothing like a bell. It is highest at 0 and trails off to the right: it is skewed to the right. Its mean is books, and its variance is , so .
The population, as percentages of students: 40% borrowed no books, 30% one, 15% two, 10% three and 5% four. The tallest bar is at 0, and the bars fall away to the right.
Means of two
Choose 2 students at random and take the mean number of books they borrowed. It can be 0, 0.5, 1, and so on up to 4. The probability that both borrowed none is 0.4 × 0.4 = 0.16; the mean is 0.5 when one borrowed none and the other one, which has probability 2 × 0.4 × 0.3 = 0.24.
The distribution of the mean of 2 is still skewed to the right, but less so than the population: the mean of 2 is 0 only 16% of the time, against 40% for one student, because both students now have to borrow nothing.
The mean of 2 students, as percentages: highest at 0.5, and still trailing off to the right.
Means of ten
Now take the mean of 10 students. Its distribution, worked out exactly from the population, is almost symmetric. It peaks at 1.0 and 1.1, right at , and falls away at about the same rate on both sides: the chance of a mean of 0.5 or less is 0.061, and the chance of a mean of 1.7 or more, the same distance above 1.1, is 0.077.
The center has not moved: the mean of the distribution is still 1.1. Its variance is 1.39 ÷ 10 = 0.139, as Var says. What has changed is the shape. The lopsided population has produced sample means that pile up in a near-symmetric hump.
The mean of 10 students, as percentages, in steps of 0.1 from 0 to 2.5. The hump is nearly symmetric about 1.1, with a slightly longer tail on the right.
The central limit theorem
This is what the central limit theorem says. Take random samples of size n from any population with mean and variance . When n is large, the sample mean X̄ is approximately normally distributed, with mean and variance . That is, X̄ is approximately , and the larger n is, the better the approximation. It holds whatever the shape of the population.
The theorem is about the means of samples, not about the population. The population of borrowings is just as skewed after any number of samples are taken; it is the distribution of the sample mean that becomes normal.
How large is large? About 30 is usually enough. A population that is already symmetric needs fewer, and one that is badly skewed needs more. If the population is itself normal, X̄ is exactly normal for every n.
the means of a skewed population pile into a symmetric bell, and their spread is σ/√n — quadrupling n halves it; here σ/√n = 0.25
Take n to 100 and read the spread
Six hundred sample means from a population with mean 1 and standard deviation 1 that is strongly skewed: most values are small and a few are large. Drag n from 16 up to 100, and the pile of means grows narrower and more symmetric, a bell centered on 1.
Why the normal curve is everywhere
Many quantities are the sum or the mean of a large number of small, independent effects. A person’s height depends on many genes and on years of food and health; an error in a measurement is the sum of many small disturbances. The central limit theorem says that such sums and means come out close to normal, whatever each small effect looks like on its own. That is why so many measurements follow a normal curve.
Using it
Take a random sample of 30 students. By the central limit theorem, their mean number of borrowings is approximately normal, with mean 1.1 and standard error .
To find the probability that the sample mean is more than 1.5, standardize: . Then . Worked out exactly from the population, the probability is 0.031, so the normal curve is close even though the population is far from normal.
The normal curve for the mean of 30 students, centered on 1.1 with standard error 0.2153. The shaded area to the right of 1.5 is about 0.032.
The usual mistakes
Using the normal curve for a single value. One student’s borrowings follow the skewed population; the theorem is about the mean of many students.
Thinking the population becomes normal. Taking samples does not change the population; only the distribution of the sample mean changes shape.
Standardizing with instead of . For the mean of 30 students the spread is 0.2153, not 1.18.
Trusting it for a small sample from a skewed population. The mean of 2 students above is still clearly skewed, and a normal curve would give it the wrong probabilities.
Worked example: Skewed Service Times at a Post Office Counter, and the Mean Time for Sixty-Four Customers
Question The time a post office clerk takes to serve a customer has mean 4 minutes and standard deviation 4 minutes. The distribution is strongly skewed: most customers take a minute or two, and a few take far longer. (a) Using Φ(1.6) = 0.9452, find the probability that the mean service time of a random sample of 64 customers is more than 4.8 minutes. (b) A trainee suggests using the same method for the mean of just 4 customers. Using Φ(2) = 0.9772, find the probability that the method would then assign to a mean service time below 0 minutes, and explain what this shows.
1.The service times are not normal, but n = 64 is large, so by the central limit theorem X is approximately normal. Its mean is μ = 4 minutes and its standard deviation is σ√n = 4√64 = 48 = 0.5 minutes.
A single service time is strongly skewed: most are short and a few are very long. 2.Standardize: z = 4.8 − 40.5 = 1.6.
By the central limit theorem, the mean of 64 customers is approximately N(4, 0.52). 3.(a) P(X > 4.8) = 1 − Φ(1.6) = 1 − 0.9452 = 0.0548.
(a) The shaded tail beyond 4.8 minutes is 1 − Φ(1.6) = 0.0548. 4.For 4 customers the standard deviation of the mean would be 4√4 = 2 minutes. A normal curve with mean 4 and standard deviation 2 puts 0 minutes at z = 0 − 42 = −2, so it gives P(X < 0) = 1 − Φ(2) = 1 − 0.9772 = 0.0228.
For 4 customers the same method uses N(4, 22), and the shaded tail lies below 0 minutes. 5.(b) The method gives a probability of 0.0228 to a negative mean time, which is impossible. For a population this skewed, 4 customers are far too few for the mean to be close to normal. With 64 customers, 0 minutes is 8 standard deviations below the mean, and the normal curve puts almost no area there.
(b) The method gives 0.0228 to a negative mean time, which is impossible: 4 customers are too few.
Answer: (a) P(X > 4.8) = 0.0548; (b) 0.0228 for a negative mean time, which is impossible, so 4 customers are too few for the central limit theorem to apply
Common mistakes
- Using the normal distribution for a single customer, P(X > 4.8) = 1 − Φ(0.2). One service time is strongly skewed, not normal; the theorem is about the mean of many customers.
- Using σ = 4 rather than σ√n = 0.5 for the mean, which gives z = 0.2 and a probability of about 0.42. The mean of 64 customers has a standard deviation 8 times smaller than one customer's time.