The Sampling Distribution

Every possible sample mean, piled up.

One sample, one mean

Here is a small population: ten scores, 3, 5, 4, 8, 2, 6, 5, 7, 4 and 6. Their total is 50, so the population mean is 50 ÷ 10 = 5.

Take a sample of four of them. One sample gets the scores 5, 8, 5 and 6. Its total is 24, so its mean is 24 ÷ 4 = 6. The sample mean, 6, is an estimate of the population mean, 5, and here it is 1 too high.

population3548265746mean 5sample 15856mean 6

The ten scores, with mean 5, and below them the four that the first sample took: 5, 8, 5 and 6, with mean 6.

Another sample, another mean

Take a second sample of four from the same ten scores. It gets 3, 2, 6 and 7, with total 18 and mean 18 ÷ 4 = 4.5. A third sample gets 4, 6, 4 and 6, with total 20 and mean 20 ÷ 4 = 5.

The population has not changed, but each sample holds different scores, so each one has a different mean: 6, then 4.5, then 5. Before a sample is taken, its mean is not known. The sample mean is a random variable, and like any random variable it has a distribution.

population3548265746mean 5sample 15856mean 6sample 23267mean 4.5sample 34646mean 5

Three samples of four from the same ten scores. Their means are 6, 4.5 and 5.

Every sample at once

The distribution of the sample mean over all the samples that could be taken is called its sampling distribution. Here it can be found completely. The number of ways to choose 4 scores from 10 is 10C4 = (10 × 9 × 8 × 7) / (4 × 3 × 2 × 1) = 5,040 ÷ 24 = 210, so there are 210 possible samples, each with its own mean.

The smallest possible mean comes from the four lowest scores, 2, 3, 4 and 4: 13 ÷ 4 = 3.25. The largest comes from the four highest, 8, 7, 6 and 6: 27 ÷ 4 = 6.75. Most of the means are much closer to 5 than that. Of the 210 samples, 186 have a mean from 4 to 6, almost 9 in every 10, while the scores themselves run all the way from 2 to 8.

Averaging pulls the means together. A sample that holds the 8 usually holds some lower scores too, and they pull its mean back toward the middle.

3.544.555.566.5

The means of all 210 samples of four, in steps of 0.25 from 3.25 to 6.75. The tallest bar is at 5, with 28 samples, and 186 of the 210 means lie from 4 to 6.

Centered on the population mean

Add up all 210 sample means and divide by 210, and the answer is exactly 5, the population mean. Single samples miss it, some too high and some too low, but over all the possible samples the misses cancel. This is the sense in which the sample mean is a fair estimate of the population mean.

So the sampling distribution has the same center as the population, and a much narrower spread. Larger samples pull the means in even more tightly.

2345678population mean 5

Fifteen sample means drawn on the scores’ own axis, from 2 to 8. They pile up around 5 and fill only the middle third of the axis.

Three distributions

Three different distributions are easy to mix up. The population distribution is the ten scores themselves, from 2 to 8. The distribution of one sample is the four scores it happened to take, such as 5, 8, 5 and 6. The sampling distribution is the distribution of the mean over every possible sample, from 3.25 to 6.75 and centered on 5.

The sampling distribution is the one that says how far a sample mean is likely to be from the population mean, which is what anyone using a sample needs to know.

The usual mistakes

Expecting the sample means to spread as widely as the scores. Each mean blends four scores, so an extreme score is diluted by the others drawn with it.

Expecting a sample mean to equal the population mean. Only 28 of the 210 samples have a mean of exactly 5; it is the average of all the sample means that equals 5.

Worked example: Cars Owned by the Five Households on a Short Street, and Every Sample of Two That Could Be Chosen

Question The five households on a short street own 0, 1, 1, 2 and 3 cars. A researcher chooses 2 of the households at random, without replacement, and works out X, the mean number of cars in her sample. (a) List all the possible samples and find the sampling distribution of X. Hence find P(X ≥ 2). (b) Find E(X), and compare it with the mean number of cars per household on the street.

  1. 1.Call the households A, B, C, D and E, with 0, 1, 1, 2 and 3 cars. The ten samples and their means are AB 0.5, AC 0.5, AD 1, AE 1.5, BC 1, BD 1.5, BE 2, CD 1.5, CE 2 and DE 2.5.

    A: 0B: 1C: 1D: 2E: 3AB 0.5AC 0.5AD 1AE 1.5BC 1BD 1.5BE 2CD 1.5CE 2DE 2.510 samples, all equally likely
    A: 0B: 1C: 1D: 2E: 3AB 0.5AC 0.5AD 1AE 1.5BC 1BD 1.5BE 2CD 1.5CE 2DE 2.510 samples, all equally likely
    There are 52 = 10 samples of two households, each with its own mean.
  2. 2.Count how often each mean occurs. X takes the values 0.5, 1, 1.5, 2 and 2.5 with probabilities 210, 210, 310, 210 and 110. Check: 2 + 2 + 3 + 2 + 1 = 10.

    2/100.52/1013/101.52/1021/102.5mean of samplemeans: 0.5, 1, 1.5, 2, 2.5counts: 2, 2, 3, 2, 1
    2/100.52/1013/101.52/1021/102.5mean of samplemeans: 0.5, 1, 1.5, 2, 2.5counts: 2, 2, 3, 2, 1
    The sampling distribution of X: each bar counts the samples with that mean.
  3. 3.(a) P(X ≥ 2) = 210 + 110 = 310, from the samples BE, CE and DE.

    2/100.52/1013/101.52/1021/102.5mean of sampleat least 2: 2/10 + 1/10 = 3/10
    2/100.52/1013/101.52/1021/102.5mean of sampleat least 2: 2/10 + 1/10 = 3/10
    (a) The samples BE, CE and DE have a mean of at least 2: P(X ≥ 2) = 310.
  4. 4.Weight each value by its probability: E(X) = 0.5 × 2 + 1 × 2 + 1.5 × 3 + 2 × 2 + 2.5 × 110 = 1410 = 1.4.

    0.511.522.5mean of samplemean 1.4(0.5 × 2 + 1 × 2 + 1.5 × 3 + 2 × 2 + 2.5)/10= 14/10 = 1.4
    0.511.522.5mean of samplemean 1.4(0.5 × 2 + 1 × 2 + 1.5 × 3 + 2 × 2 + 2.5)/10= 14/10 = 1.4
    Weighting each mean by its probability gives E(X) = 1.4.
  5. 5.(b) The mean for the street is μ = 0 + 1 + 1 + 2 + 35 = 75 = 1.4 cars, so E(X) = μ. The sample mean is an unbiased estimator of the population mean, even though no single sample has a mean of exactly 1.4.

    0.511.522.5mean of samplemean 1.4street: (0 + 1 + 1 + 2 + 3)/5 = 1.4mean of the means = street mean
    0.511.522.5mean of samplemean 1.4street: (0 + 1 + 1 + 2 + 3)/5 = 1.4mean of the means = street mean
    (b) E(X) = μ = 1.4, although no single sample has a mean of 1.4.

Answer: (a) X is 0.5, 1, 1.5, 2 or 2.5 with probabilities 210, 210, 310, 210 and 110, and P(X ≥ 2) = 310; (b) E(X) = 1.4, the same as the mean for the street, 1.4 cars per household

Common mistakes

  • Treating the two households with one car as one household and listing fewer samples. B and C are different households, so AB and AC are two different samples, and leaving one out gives the wrong probabilities.
  • Expecting a sample mean to equal the population mean. Each sample mean is 0.5, 1, 1.5, 2 or 2.5; it is the average over all the possible samples that equals 1.4.

More sampling problems, worked step by step →

Practice The Sampling Distribution in the app