Two lines, two spreads
Two production lines make the same part. Their mean sizes are much the same, but line B’s parts look more variable than line A’s. The question is whether the populations differ in spread, against , not in mean.
The chi-squared test of the last lesson compares one variance with a claimed number. Here neither variance is known: both are estimated from samples, and the two estimates are compared with each other.
A ratio, not a difference
Two spreads are compared by dividing. If both populations are normal with the same variance, the ratio of the two sample variances follows an F distribution, which depends on two numbers of degrees of freedom: one for the variance on top and one for the variance underneath.
By convention the larger sample variance goes on top: F = larger smaller . Then F is at least 1, and only values far above 1 count against equal spreads.
The two samples
Line B: 10 parts, with . Line A: 8 parts, with . Each divides the squared deviations from its own sample mean by one less than its sample size.
The larger is 9.6, so . Each sample spends one degree of freedom on its own mean: line B, on top, has 10 − 1 = 9, and line A, underneath, has 8 − 1 = 7. F is read on F(9, 7), top first.
The gold curve is the F distribution with 9 and 7 degrees of freedom, highest at about 0.6. The shaded tail from 3.68 holds 5 percent of the area. The lines’ F = 2.4 is to its left.
The verdict
With 9 and 7 degrees of freedom, , so the tail begins at 3.68. 2.4 < 3.68, so do not reject : . The p-value is .
Stated about the lines: there is no evidence that their spreads differ. Line B’s sample variance is 2.4 times line A’s, but samples of 10 and 8 parts from equally variable lines give ratios that large about 13 times in 100.
Which level 3.68 tests at
Suppose line B was rebuilt, and the question set before measuring was whether it is now less consistent: : . Then all 5 percent goes in the upper tail, and 3.68 is the 5 percent critical value.
Suppose instead the question was only whether the spreads differ: : . Either line could have come out on top. Putting the larger variance on top turns a ratio far below 1 into one far above 1, so the lower tail is never used, but each line had a 5 percent chance of landing in the upper tail. Comparing with 3.68 is then a test at 10 percent.
For that question at 5 percent, use the upper 2.5 percent point instead: for 9 and 7 degrees of freedom. 2.4 is short of both 3.68 and 4.82, so the verdict here is the same at every one of these levels.
Tests for means and tests for spreads
A t-test asks whether two means differ, and an F-test asks whether two spreads differ.
The pooled two-sample t-test assumes the two populations have equal variances. An F-test that finds no evidence of a difference leaves that assumption standing; one that rejects says the pooled test is the wrong tool for those samples.
Like the chi-squared test for one variance, the F-test assumes both populations are normal, and it is badly misled when they are not.
The usual mistakes
Putting the smaller variance on top. 4.0 ÷ 9.6 = 0.42 is a ratio below 1, and the upper-tail critical value cannot judge it.
Subtracting. 9.6 − 4.0 = 5.6 follows no F distribution; spreads are compared by their ratio.
Swapping the degrees of freedom. F(7, 9) is a different curve, with upper 5 percent point 3.293, not 3.677. The first number belongs to the variance on top.
Using standard deviations. is not F; F is a ratio of variances.
Bolts from two machines
In the application below, part (b) compares machine A, mm² from 11 bolts, with machine B, mm² from 13 bolts. The larger estimate goes on top, giving F = 2.5 on 10 and 12 degrees of freedom. Part (a) is the chi-squared test of the last lesson.
Worked example: Bolts from Two Machines, One Tested Against the Specification's Variance and Then Against the Other Machine
Question A factory's specification says that the lengths of its bolts must have a variance of no more than 0.04 mm2. A sample of 11 bolts from machine A gives an unbiased variance estimate sA2 = 0.08 mm2, and a sample of 13 bolts from machine B gives sB2 = 0.032 mm2. Lengths are normally distributed. (a) Test at the 5% level whether the variance of machine A is greater than 0.04 mm2. (b) Test at the 5% level whether machine A's lengths vary more than machine B's.
1.For machine A, H0: σA2 = 0.04 and H1: σA2 > 0.04. The test statistic is (n − 1)sA20.04 = 10 × 0.080.04 = 20.0, on the χ2(10) distribution.
For machine A the statistic (n − 1)s2σ2 = 20.0 follows χ2(10) if H0 is true. 2.(a) The upper 5% critical value of χ2(10) is 18.307. Since 20.0 > 18.307, reject H0: there is evidence at the 5% level that machine A's variance is greater than the 0.04 mm2 the specification allows.
(a) The top 5% of χ2(10) lies beyond 18.307, and 20.0 is in it: reject H0. 3.For the comparison, H0: σA2 = σB2 and H1: σA2 > σB2. The test statistic is the ratio of the estimates, the larger on top: F = 0.080.032 = 2.5.
The ratio of the two variance estimates, the larger on top, is F = 2.5. 4.The degrees of freedom are 11 − 1 = 10 for the top and 13 − 1 = 12 for the bottom. The upper 5% critical value of F(10, 12) is 2.753.
The top 5% of F(10, 12) lies beyond 2.753. 5.(b) 2.5 < 2.753, so do not reject H0. There is not enough evidence at the 5% level that machine A's lengths vary more than machine B's. The two results are consistent: A is shown to be outside the specification, but samples of 11 and 13 bolts are too small to show that A is worse than B.
(b) 2.5 is just outside the critical region, so do not reject H0.
Answer: (a) the test statistic is 20.0, above the critical value 18.307 of χ2(10), so reject H0: machine A's variance is greater than 0.04 mm2; (b) F = 2.5, below the critical value 2.753 of F(10, 12), so do not reject H0: there is not enough evidence that A's lengths vary more than B's
Common mistakes
- Using the standard deviations in place of the variances: 10 × √0.080.2 or F = √0.08√0.032 = 1.58. Both statistics are built from variances, s2 and σ2.
- Swapping the degrees of freedom and reading F(12, 10), whose critical value is 2.913. The first number belongs to the estimate on top, machine A's, with 10 degrees of freedom.