The population and the sample
A school has 1,000 students, and you want to know how many hours a night they sleep. The whole group you want to know about, all 1,000 students, is the population.
Asking all 1,000 takes a long time, so you ask 30 of them. The part of the population you actually measure is the sample. Here the sample is 30 students out of 1,000, which is 3% of the population.
The rectangle is the population of 1,000 students. The circle is the sample of 30, and the other 970 students are not asked.
Parameters and statistics
A number that describes the whole population is a parameter. The mean number of hours slept by all 1,000 students is a parameter. It is one fixed number, but you do not know it: the only way to find it exactly is to measure every student.
A number worked out from the sample is a statistic. The mean of the 30 students you asked is a statistic, and you can calculate it, because you have their answers.
The usual symbols keep the two apart. The population mean is and the population standard deviation is ; these are parameters. The sample mean is x̄ and the sample standard deviation is s; these are statistics.
Proportions work the same way. The true proportion of faulty parts made by a factory is a parameter. The proportion faulty in the one box you opened is a statistic.
A statistic changes from sample to sample
To see the difference, take a population small enough to list: 10 students, who sleep 6, 8, 7, 9, 5, 8, 7, 6, 9 and 7 hours. Their total is 72 hours, so the population mean is hours.
Now take samples of 4. One sample gets the students who sleep 6, 8, 9 and 7 hours, with mean 30 ÷ 4 = 7.5 hours. Another gets 7, 5, 6 and 7 hours, with mean 25 ÷ 4 = 6.25 hours. A third gets 8, 8, 9 and 7 hours, with mean 32 ÷ 4 = 8 hours.
The population mean is 7.2 every time, because the population has not changed. The sample mean is different for each sample, because each sample holds different students. None of the three is exactly 7.2, and each one is an estimate of it.
So a statistic is used as an estimate of a parameter, and its value depends on which members happen to be in the sample. The difference between a statistic and the parameter it estimates is called sampling error. It is not a mistake by the person who did the measuring: it comes from measuring only part of the population.
The top row is the whole population, with mean 7.2 hours. Each row below takes 4 of the same 10 students, and the three sample means are 7.5, 6.25 and 8.
Why take a sample at all
Measuring the whole population, a census, gives the parameter exactly. A sample is used when a census costs too much, takes too long, or is impossible. To find how long a batch of light bulbs lasts, each bulb tested has to burn out, so testing every bulb would leave none to sell.
A sample is useful only when it is like the population. A sample of 30 students chosen at random from the school register is likely to be. A sample of the 30 students in the library at 8 a.m. is not: early arrivals may sleep less than everyone else, and asking more of them would not fix that. How to choose a sample is the subject of Sampling Methods.
From the sample to the population
Suppose 12 of the 30 students in a random sample sleep less than 7 hours. The sample proportion is 12 ÷ 30 = 0.4, a statistic. It estimates the proportion of the whole school who sleep less than 7 hours, so about 0.4 × 1,000 = 400 students.
The 400 is an estimate, not a count. A different sample of 30 would have given a different proportion, and a different estimate.
The usual mistakes
Calling a number from a sample a parameter. The mean height of 40 people you measured is a statistic, however carefully they were measured; the mean height of everyone in the country is the parameter.
Expecting the sample mean to equal the population mean. It is an estimate, and a second sample would give a different value.
Treating the parameter as something that changes. The population mean is fixed; only the estimates of it change.