The Two-Sample T-Test

Comparing two means with the spread unknown.

Two groups of plants

Sixteen seedlings of the same kind are split at random into two groups of 8. One group gets plain water, the other water with plant food. After six weeks their heights, in centimeters, are:

plain: 16, 17, 18, 20, 20, 21, 23, 25, with mean 160 ÷ 8 = 20 cm;

fed: 19, 22, 22, 23, 25, 26, 27, 28, with mean 192 ÷ 8 = 24 cm.

The fed plants are 4 cm taller on average. But the two groups overlap a lot, and two groups of plain-watered plants would also have different means. The question is whether 4 cm is more than that kind of chance would give.

16171819202122232425262728

The eight plants given plain water, one dot each. Their heights run from 16 cm to 25 cm, with mean 20 cm.

16171819202122232425262728

The eight plants given plant food, on the same scale. Their heights run from 19 cm to 28 cm, with mean 24 cm.

The hypotheses

Let μ₁ be the mean height of all such seedlings grown with plain water, and μ₂ the mean with plant food. H₀: μ₁ = μ₂. Under H₀, both groups come from populations with one shared mean, and the 4 cm gap is sampling variation.

The question asked before the trial was whether the food makes the plants taller, so H₁: μ₂ > μ₁, a one-tailed test at 5 percent. Had the question been only whether the food makes a difference, H₁ would be μ₁ ≠ μ₂, a two-tailed test.

Spread from the samples

A z-test would divide the gap by a standard error built from the population standard deviation. Nobody knows it for these seedlings, so the spread has to be estimated from the two samples, as s was in the one-sample t-test.

For the plain group, the deviations from 20 are −4, −3, −2, 0, 0, 1, 3 and 5. Their squares add up to 16 + 9 + 4 + 0 + 0 + 1 + 9 + 25 = 64, so s₁² = 64/7 = 9.143.

For the fed group, the deviations from 24 are −5, −2, −2, −1, 1, 2, 3 and 4. Their squares add up to 25 + 4 + 4 + 1 + 1 + 4 + 9 + 16 = 64, so s₂² = 64/7 = 9.143 as well.

One pooled variance

The pooled test assumes the two populations have the same variance, and combines the two estimates into one, each weighted by its degrees of freedom: pooled s² = ((n₁ − 1)s₁² + (n₂ − 1)s₂²)/(n₁ + n₂ − 2) = (64 + 64)/14 = 9.143.

Each sample spends one degree of freedom on its own mean, so the pooled estimate has n₁ + n₂ − 2 = 8 + 8 − 2 = 14 degrees of freedom, and so does the test.

The gap in standard errors

The standard error of the gap between the two sample means is √(pooled s² × (1/n₁ + 1/n₂)) = √(9.143 × (1/8 + 1/8)) = √2.286 = 1.512 cm.

So t = (x̄₂ − x̄₁)/1.512 = (24 − 20)/1.512 = 2.646. The fed mean is 2.646 standard errors above the plain mean.

t2.646

The t distribution with 14 degrees of freedom. The shaded tail beyond the plants’ t = 2.646 is the p-value, 0.0096: under one percent of the area.

From t to p

The p-value is the chance, if H₀ is true, of a t at least this large: P(T ≥ 2.646) = 0.0096 on 14 degrees of freedom, about 0.01. A calculator gives it directly from the two lists.

0.0096 < 0.05, so reject H₀. Equally, the one-tailed 5 percent critical value is 1.761, and 2.646 is beyond it.

Stated about the plants: there is evidence at the 5 percent level that seedlings given plant food grow taller, on average, than seedlings given plain water.

With the two-tailed question, μ₁ ≠ μ₂, the p-value doubles to 0.0192 and the critical value is 2.145. H₀ is rejected at 5 percent either way.

What the test assumes

The samples must be independent: each plant is in one group only, and one plant’s height says nothing about another’s. The same plants measured before and after feeding would not be two independent samples.

Each population should be roughly normal. With only eight plants a group, that is assumed, not checked.

The pooled form also assumes equal population variances. Here both sample variances are 9.143, so nothing argues against it. When the two sample variances are very different, the F-test, two lessons on, tests whether the populations’ spreads differ.

The sample sizes do not need to be equal. With n₁ and n₂ different, the same formulas hold, and the degrees of freedom are still n₁ + n₂ − 2.

The usual mistakes

Using one sample’s degrees of freedom. 8 − 1 = 7 belongs to a one-sample test; the pooled estimate uses both samples, which gives 14.

Adding the two variances without weighting or dividing. 9.143 + 9.143 = 18.29 is not the pooled variance; the pooled variance is their weighted average.

Dividing the gap by the pooled standard deviation instead of by the standard error of the gap. √9.143 = 3.02 cm is the spread of single plants; the gap between two means of 8 varies less.

Saying the means are proved equal when H₀ is not rejected. The test only finds whether the data give enough evidence against H₀.

Tomato seedlings and two fertilizers

In the application below, 6 seedlings get fertilizer A and 6 get fertilizer B. The two variance estimates are pooled, the test is two-tailed on 6 + 6 − 2 = 10 degrees of freedom, and the verdict is taken at 5 percent and at 1 percent.

Worked example: Tomato Seedlings Grown with Two Fertilizers and Compared by Their Mean Heights

Question A gardener grows 12 tomato seedlings, giving fertilizer A to 6 of them and fertilizer B to the other 6, chosen at random. After four weeks the heights in centimeters are: A: 22, 22, 24, 25, 25, 26; B: 19, 19, 21, 21, 23, 23. Heights are assumed to be normal, with the same variance for both fertilizers. (a) Find the pooled estimate of the variance and the test statistic for comparing the two means. (b) Test at the 5% level whether the fertilizers give different mean heights, and then at the 1% level.

  1. 1.Let μA and μB be the mean heights. H0: μA = μB and H1: μA ≠ μB, a two-tailed test. The sample means are xA = 1446 = 24 cm and xB = 1266 = 21 cm.

    1820222426ABmean 24mean 21A: mean 144/6 = 24 cmB: mean 126/6 = 21 cm
    1820222426ABmean 24mean 21A: mean 144/6 = 24 cmB: mean 126/6 = 21 cm
    The heights of the seedlings, in centimeters, with each sample's mean.
  2. 2.The squared deviations add up to 4 + 4 + 0 + 1 + 1 + 4 = 14 for A and 4 + 4 + 0 + 0 + 4 + 4 = 16 for B, so sA2 = 145 = 2.8 and sB2 = 165 = 3.2.

    1820222426ABmean 24mean 21A: variance 14/5 = 2.8B: variance 16/5 = 3.2
    1820222426ABmean 24mean 21A: variance 14/5 = 2.8B: variance 16/5 = 3.2
    Each sample's unbiased variance divides its squared deviations by 6 − 1 = 5.
  3. 3.The pooled variance weights each estimate by its degrees of freedom: sp2 = 5 × 2.8 + 5 × 3.210 = 3010 = 3.

    1820222426ABmean 24mean 21pooled: (5 × 2.8 + 5 × 3.2)/10 = 3
    1820222426ABmean 24mean 21pooled: (5 × 2.8 + 5 × 3.2)/10 = 3
    The two estimates, each with 5 degrees of freedom, pool to sp2 = 3.
  4. 4.(a) The standard error of the difference is √3 × (16 + 16) = √1 = 1 cm, so the test statistic is 24 − 211 = 3.0, on 6 + 6 − 2 = 10 degrees of freedom.

    −2.2282.2280test statistic 3.010 df, if H0 is true(24 − 21)/√(3 × 2/6) = 3.0on 6 + 6 − 2 = 10 degrees of freedom
    −2.2282.2280test statistic 3.010 df, if H0 is true(24 − 21)/√(3 × 2/6) = 3.0on 6 + 6 − 2 = 10 degrees of freedom
    (a) The test statistic is 3.0 on t(10). The shaded tails beyond ± 2.228 are the 5% critical region, and 3.0 is in it.
  5. 5.(b) At 5%, two-tailed, the critical value of t(10) is 2.228, and 3.0 > 2.228: reject H0. There is evidence that the fertilizers give different mean heights, with the seedlings given A taller. At 1% the critical value is 3.169, and 3.0 < 3.169, so at that level the difference is not significant.

    −3.1693.1690test statistic 3.010 df, if H0 is true3.0 > 2.228: reject H0 at 5%3.0 < 3.169: not at 1%
    −3.1693.1690test statistic 3.010 df, if H0 is true3.0 > 2.228: reject H0 at 5%3.0 < 3.169: not at 1%
    (b) At 1% the tails lie beyond ± 3.169, and 3.0 falls just short of them.

Answer: (a) the pooled variance is 3 and the test statistic is 3.0, on 10 degrees of freedom; (b) 3.0 > 2.228, so at 5% reject H0: the fertilizers give different mean heights; at 1% the critical value is 3.169, so there the difference is not significant

Common mistakes

  • Using 6 − 1 = 5 degrees of freedom, as for one sample. The pooled estimate uses both samples, which gives 5 + 5 = 10 degrees of freedom.
  • Adding the two variances without weighting or dividing, 2.8 + 3.2 = 6, and using √66. The pooled variance is the weighted average 3, and it is multiplied by 16 + 16.

More statistical inference problems, worked step by step →

Practice The Two-Sample T-Test in the app