Comparing Two Variances with the F-Test

One spread over the other, larger on top.

Two lines, two spreads

Two production lines make the same part. Their mean sizes are much the same, but line B’s parts look more variable than line A’s. The question is whether the populations differ in spread, σ_B² against σ_A², not in mean.

The chi-squared test of the last lesson compares one variance with a claimed number. Here neither variance is known: both are estimated from samples, and the two estimates are compared with each other.

A ratio, not a difference

Two spreads are compared by dividing. If both populations are normal with the same variance, the ratio of the two sample variances follows an F distribution, which depends on two numbers of degrees of freedom: one for the variance on top and one for the variance underneath.

By convention the larger sample variance goes on top: F = larger s² ÷ smaller s². Then F is at least 1, and only values far above 1 count against equal spreads.

The two samples

Line B: 10 parts, with s_B² = 9.6. Line A: 8 parts, with s_A² = 4.0. Each s² divides the squared deviations from its own sample mean by one less than its sample size.

The larger is 9.6, so F = 9.6 / 4.0 = 2.4. Each sample spends one degree of freedom on its own mean: line B, on top, has 10 − 1 = 9, and line A, underneath, has 8 − 1 = 7. F is read on F(9, 7), top first.

F2.43.68

The gold curve is the F distribution with 9 and 7 degrees of freedom, highest at about 0.6. The shaded tail from 3.68 holds 5 percent of the area. The lines’ F = 2.4 is to its left.

The verdict

With 9 and 7 degrees of freedom, P(F ≥ 3.677) = 0.05, so the tail begins at 3.68. 2.4 < 3.68, so do not reject H₀: σ_A² = σ_B². The p-value is P(F ≥ 2.4) = 0.1307.

Stated about the lines: there is no evidence that their spreads differ. Line B’s sample variance is 2.4 times line A’s, but samples of 10 and 8 parts from equally variable lines give ratios that large about 13 times in 100.

Which level 3.68 tests at

Suppose line B was rebuilt, and the question set before measuring was whether it is now less consistent: H₁: σ_B² > σ_A². Then all 5 percent goes in the upper tail, and 3.68 is the 5 percent critical value.

Suppose instead the question was only whether the spreads differ: H₁: σ_A² ≠ σ_B². Either line could have come out on top. Putting the larger variance on top turns a ratio far below 1 into one far above 1, so the lower tail is never used, but each line had a 5 percent chance of landing in the upper tail. Comparing with 3.68 is then a test at 10 percent.

For that question at 5 percent, use the upper 2.5 percent point instead: P(F ≥ 4.823) = 0.025 for 9 and 7 degrees of freedom. 2.4 is short of both 3.68 and 4.82, so the verdict here is the same at every one of these levels.

Tests for means and tests for spreads

A t-test asks whether two means differ, and an F-test asks whether two spreads differ.

The pooled two-sample t-test assumes the two populations have equal variances. An F-test that finds no evidence of a difference leaves that assumption standing; one that rejects H₀ says the pooled test is the wrong tool for those samples.

Like the chi-squared test for one variance, the F-test assumes both populations are normal, and it is badly misled when they are not.

The usual mistakes

Putting the smaller variance on top. 4.0 ÷ 9.6 = 0.42 is a ratio below 1, and the upper-tail critical value cannot judge it.

Subtracting. 9.6 − 4.0 = 5.6 follows no F distribution; spreads are compared by their ratio.

Swapping the degrees of freedom. F(7, 9) is a different curve, with upper 5 percent point 3.293, not 3.677. The first number belongs to the variance on top.

Using standard deviations. √9.6 ÷ √4.0 = 1.55 is not F; F is a ratio of variances.

Bolts from two machines

In the application below, part (b) compares machine A, s² = 0.08 mm² from 11 bolts, with machine B, s² = 0.032 mm² from 13 bolts. The larger estimate goes on top, giving F = 2.5 on 10 and 12 degrees of freedom. Part (a) is the chi-squared test of the last lesson.

Worked example: Bolts from Two Machines, One Tested Against the Specification's Variance and Then Against the Other Machine

Question A factory's specification says that the lengths of its bolts must have a variance of no more than 0.04 mm2. A sample of 11 bolts from machine A gives an unbiased variance estimate sA2 = 0.08 mm2, and a sample of 13 bolts from machine B gives sB2 = 0.032 mm2. Lengths are normally distributed. (a) Test at the 5% level whether the variance of machine A is greater than 0.04 mm2. (b) Test at the 5% level whether machine A's lengths vary more than machine B's.

  1. 1.For machine A, H0: σA2 = 0.04 and H1: σA2 > 0.04. The test statistic is (n − 1)sA20.04 = 10 × 0.080.04 = 20.0, on the χ2(10) distribution.

    0test statistic 20.0chi-sq(10) if H0 is trueH1: variance of A > 0.0410 × 0.08/0.04 = 20.0 on chi-sq(10)
    0test statistic 20.0chi-sq(10) if H0 is trueH1: variance of A > 0.0410 × 0.08/0.04 = 20.0 on chi-sq(10)
    For machine A the statistic (n − 1)s2σ2 = 20.0 follows χ2(10) if H0 is true.
  2. 2.(a) The upper 5% critical value of χ2(10) is 18.307. Since 20.0 > 18.307, reject H0: there is evidence at the 5% level that machine A's variance is greater than the 0.04 mm2 the specification allows.

    18.3070test statistic 20.0chi-sq(10) if H0 is true20.0 > 18.307: reject H0A varies more than the specification
    18.3070test statistic 20.0chi-sq(10) if H0 is true20.0 > 18.307: reject H0A varies more than the specification
    (a) The top 5% of χ2(10) lies beyond 18.307, and 20.0 is in it: reject H0.
  3. 3.For the comparison, H0: σA2 = σB2 and H1: σA2 > σB2. The test statistic is the ratio of the estimates, the larger on top: F = 0.080.032 = 2.5.

    01F = 2.5F(10, 12) if H0 is trueH1: A varies more than BF = 0.08/0.032 = 2.5
    01F = 2.5F(10, 12) if H0 is trueH1: A varies more than BF = 0.08/0.032 = 2.5
    The ratio of the two variance estimates, the larger on top, is F = 2.5.
  4. 4.The degrees of freedom are 11 − 1 = 10 for the top and 13 − 1 = 12 for the bottom. The upper 5% critical value of F(10, 12) is 2.753.

    2.75301F = 2.5F(10, 12) if H0 is true10 and 12 degrees of freedomF(10, 12) at 5%: 2.753
    2.75301F = 2.5F(10, 12) if H0 is true10 and 12 degrees of freedomF(10, 12) at 5%: 2.753
    The top 5% of F(10, 12) lies beyond 2.753.
  5. 5.(b) 2.5 < 2.753, so do not reject H0. There is not enough evidence at the 5% level that machine A's lengths vary more than machine B's. The two results are consistent: A is shown to be outside the specification, but samples of 11 and 13 bolts are too small to show that A is worse than B.

    2.75301F = 2.5F(10, 12) if H0 is true2.5 < 2.753: do not reject H0A is not shown to vary more than B
    2.75301F = 2.5F(10, 12) if H0 is true2.5 < 2.753: do not reject H0A is not shown to vary more than B
    (b) 2.5 is just outside the critical region, so do not reject H0.

Answer: (a) the test statistic is 20.0, above the critical value 18.307 of χ2(10), so reject H0: machine A's variance is greater than 0.04 mm2; (b) F = 2.5, below the critical value 2.753 of F(10, 12), so do not reject H0: there is not enough evidence that A's lengths vary more than B's

Common mistakes

  • Using the standard deviations in place of the variances: 10 × √0.080.2 or F = √0.08√0.032 = 1.58. Both statistics are built from variances, s2 and σ2.
  • Swapping the degrees of freedom and reading F(12, 10), whose critical value is 2.913. The first number belongs to the estimate on top, machine A's, with 10 degrees of freedom.

More statistical inference problems, worked step by step →

Practice Comparing Two Variances with the F-Test in the app