Dividing by n comes out too small
A sample of five values is 3, 5, 6, 9 and 12. Their mean is , their deviations from it are −4, −2, −1, 2 and 5, and the squared deviations are 16, 4, 1, 4 and 25, which add to 50. Dividing by n = 5 gives 10.
That figure is the variance of these five values about their own mean. As an estimate of the population variance it is too small on average, because measures spread about the population mean , and the sample mean is not .
The sum of squared deviations is smallest when it is measured from the sample’s own mean. Measured from 6 instead, the squares are 9, 1, 0, 9 and 36, which add to 55; from 8 they also add to 55. In general the sum about any center a is , so every center other than x̄ = 7 gives more. Wherever really is, the squares measured from x̄ come out no larger than those measured from , and usually smaller.
The sum of the squared deviations of 3, 5, 6, 9 and 12 from a center a. It is least, 50, when a is the sample mean 7, and larger from any other center, such as 55 from 6.
Dividing by n − 1
Dividing the same sum by n − 1 instead of n makes the estimate a little larger, and exactly enough: the result is an unbiased estimate of , right on average over all samples. It is written .
For the sample above, , against .
Checked on every sample
Take the population 1, 3 and 5. Its mean is 3 and its variance is . Draw a sample of two, with replacement: there are nine equally likely samples, from (1, 1) to (5, 5).
A sample of two values a and b has mean , and the sum of its squared deviations is . Dividing by n = 2 gives : that is 0 for the three samples with equal values, 1 for the four samples whose values differ by 2, and 4 for the two samples 1 and 5. Its average over the nine samples is , only half of .
Dividing by n − 1 = 1 instead gives 0, 2 and 8, and the average is , exactly . Dividing by n is short by the factor , here , and dividing by n − 1 undoes it.
Dividing by n: the nine samples give three 0s, four 1s and two 4s, which balance at , below .
Dividing by n − 1: the same nine samples give three 0s, four 2s and two 8s, which balance at , exactly .
Only n − 1 deviations are free
Deviations from the sample mean always add to 0: −4 − 2 − 1 + 2 + 5 = 0. So once four of them are known, the fifth is forced. Given −4, −2, −1 and 2, the last must be 5.
Five values, but only four independent pieces of information about the spread, because one was used up in finding the mean. The number of free deviations, n − 1, is called the degrees of freedom, and it is what the sum of squares is divided by.
From the sums
The sum of squared deviations can be found from and without listing the deviations: . For 3, 5, 6, 9 and 12, and , so the sum is 50 and , as before.
The mean is still found by dividing by n: . Only the variance divides by n − 1.
The usual mistakes
Dividing by n. is the spread of these five values about their own mean, and as an estimate of it is too small on average.
Dividing by n + 1. That makes the estimate smaller still, when it was already too small; the divisor must go down, to n − 1.
Dividing the mean by n − 1 as well. ; only the sum of squared deviations is divided by n − 1.
Worked example: Strawberries Weighed from a Day's Picking, and the Mean and Spread of the Whole Crop Estimated from Two Sums
Question A fruit farm picks 25 strawberries at random from a day's crop and weighs them. The masses x grams give ∑ x = 605 and ∑ x2 = 14941. (a) Find unbiased estimates of the mean and the variance of the mass of a strawberry in the crop. (b) A worker divides the sum of squares about the mean by 25 instead of 24. Find the value this gives, say why it tends to be too small, and use your unbiased estimate from (a) to find the standard error of the sample mean.
1.The unbiased estimate of the population mean is the sample mean: x = 60525 = 24.2 g.
The 25 masses add up to 605 g, so the sample mean is x = 24.2 g. 2.The sum of squares about the mean is Sxx = ∑ x2 − (∑ x)2n = 14941 − 605225 = 14941 − 14641 = 300.
The sum of squares about the mean is ∑ x2 − (∑ x)2n = 14941 − 14641 = 300. 3.(a) The unbiased estimate of the variance divides by n − 1 = 24: s2 = 30024 = 12.5. The estimates are 24.2 g for the mean and 12.5 g2 for the variance.
(a) Shared among n − 1 = 24, the 300 gives s2 = 12.5, the unbiased estimate of the variance. 4.Dividing by 25 gives 30025 = 12. The deviations are measured from x, the center of this sample itself, so their squares are smaller on average than squares measured from the true mean μ. Dividing by n − 1 instead of n makes up for that.
Shared among 25 it gives 12, which is too small on average, because each deviation is measured from x rather than from μ. 5.(b) The worker's value is 12 g2, which tends to underestimate the variance. The standard error of the mean is estimated as √s2n = √12.525 = √0.5 = 0.707 g. Check: 24 × 12.5 = 25 × 12 = 300.
(b) The standard error of the mean is √12.525 = 0.707 g.
Answer: (a) the mean 24.2 g and the variance 12.5 g2; (b) 12 g2, too small because the deviations are measured from the sample's own mean; the standard error is about 0.707 g
Common mistakes
- Working out ∑ x2n − x2 = 597.64 − 585.64 = 12 and calling it the unbiased estimate. That formula divides by n; the unbiased estimate is nn − 1 times it, 2524 × 12 = 12.5.
- Subtracting 6052 itself from 14941, or subtracting 60525 without squaring it. The square of the total must be divided by n: 605225 = 14641 is 25 times the square of the mean.