A claim about spread
A machine cuts rods to length. Its maker claims that the lengths have a variance of mm², a standard deviation of 2 mm. A buyer suspects the lengths vary more than that.
The hypotheses are about , the variance of every rod the machine cuts: : and : . The mean length is not in question. A machine can cut to the right length on average and still be too inconsistent to use.
Scaling the sample variance
Fifteen rods are measured. Their variance is estimated by , the sum of the squared deviations from x̄ divided by n − 1 = 14. That is the sample’s figure; is the population’s, the one makes a claim about.
If the lengths are normally distributed and is true, the statistic follows a chi-squared distribution with n − 1 degrees of freedom. One degree of freedom is spent because the deviations are measured from x̄, which is worked out from the same rods.
is the sum of the squared deviations, so the statistic is that sum measured in units of the claimed . On average equals , so when is true the statistic averages n − 1 = 14. A statistic far above 14 suggests the spread is larger than claimed.
Fifteen rods
The fifteen rods give mm², so mm.
The statistic is . The claimed variance goes underneath and the sample variance on top.
The gold curve is the chi-squared distribution with 14 degrees of freedom. It is highest at 12, leans to the right, and has mean 14. The shaded tail from 23.68 holds 5 percent of the area; the rods’ 22.4 falls just short of it.
The verdict
With 14 degrees of freedom, P(χ² , so at the 5 percent level the critical region is χ² .
22.4 < 23.68, so do not reject . The p-value is P(χ² , a little above 0.05. Stated about the machine: there is not enough evidence at the 5 percent level that the lengths vary more than the claimed 4 mm².
That does not prove . A sample variance of 6.4 from fifteen rods is suspicious, but not enough by itself at this level.
Two tails, two different bars
If the buyer had asked only whether the variance differs from 4, either way, would be , and 2.5 percent would go in each tail.
The chi-squared curve is not symmetric, so the two critical values are not a number and its negative. With 14 degrees of freedom, P(χ² and P(χ² . is rejected if the statistic is below 5.629 or above 26.119.
For the fifteen rods, 22.4 lies between them, so is not rejected. A sample with mm² would give 14 × 1.5 ÷ 4 = 5.25, below 5.629: evidence that the lengths vary less than claimed.
The same curve with 95 percent of its area shaded, from 5.63 to 26.12. The 2.5 percent left of 5.63 and the 2.5 percent right of 26.12 are the two critical regions, and they are not the same distance from the peak.
The usual mistakes
Using s instead of . 14 × 2.53 ÷ 4 = 8.9 mixes a standard deviation with a variance; the statistic is built from and .
Multiplying by n. 15 × 6.4 ÷ 4 = 24 uses n where n − 1 belongs.
Dividing the wrong way. 14 × 4 ÷ 6.4 = 8.75 puts the claimed variance on top; the sample variance goes on top.
Using one critical value for a two-tailed test, or the same distance either side of the peak. Each tail has its own value: 5.629 and 26.119 here.
Bolts from two machines
In the application below, part (a) tests machine A’s variance against a specification of 0.04 mm² from 11 bolts, on 10 degrees of freedom. Part (b) compares machine A with machine B, which is the next lesson’s F-test.
Worked example: Bolts from Two Machines, One Tested Against the Specification's Variance and Then Against the Other Machine
Question A factory's specification says that the lengths of its bolts must have a variance of no more than 0.04 mm2. A sample of 11 bolts from machine A gives an unbiased variance estimate sA2 = 0.08 mm2, and a sample of 13 bolts from machine B gives sB2 = 0.032 mm2. Lengths are normally distributed. (a) Test at the 5% level whether the variance of machine A is greater than 0.04 mm2. (b) Test at the 5% level whether machine A's lengths vary more than machine B's.
1.For machine A, H0: σA2 = 0.04 and H1: σA2 > 0.04. The test statistic is (n − 1)sA20.04 = 10 × 0.080.04 = 20.0, on the χ2(10) distribution.
For machine A the statistic (n − 1)s2σ2 = 20.0 follows χ2(10) if H0 is true. 2.(a) The upper 5% critical value of χ2(10) is 18.307. Since 20.0 > 18.307, reject H0: there is evidence at the 5% level that machine A's variance is greater than the 0.04 mm2 the specification allows.
(a) The top 5% of χ2(10) lies beyond 18.307, and 20.0 is in it: reject H0. 3.For the comparison, H0: σA2 = σB2 and H1: σA2 > σB2. The test statistic is the ratio of the estimates, the larger on top: F = 0.080.032 = 2.5.
The ratio of the two variance estimates, the larger on top, is F = 2.5. 4.The degrees of freedom are 11 − 1 = 10 for the top and 13 − 1 = 12 for the bottom. The upper 5% critical value of F(10, 12) is 2.753.
The top 5% of F(10, 12) lies beyond 2.753. 5.(b) 2.5 < 2.753, so do not reject H0. There is not enough evidence at the 5% level that machine A's lengths vary more than machine B's. The two results are consistent: A is shown to be outside the specification, but samples of 11 and 13 bolts are too small to show that A is worse than B.
(b) 2.5 is just outside the critical region, so do not reject H0.
Answer: (a) the test statistic is 20.0, above the critical value 18.307 of χ2(10), so reject H0: machine A's variance is greater than 0.04 mm2; (b) F = 2.5, below the critical value 2.753 of F(10, 12), so do not reject H0: there is not enough evidence that A's lengths vary more than B's
Common mistakes
- Using the standard deviations in place of the variances: 10 × √0.080.2 or F = √0.08√0.032 = 1.58. Both statistics are built from variances, s2 and σ2.
- Swapping the degrees of freedom and reading F(12, 10), whose critical value is 2.913. The first number belongs to the estimate on top, machine A's, with 10 degrees of freedom.