The One-Sample T-Test

One claimed mean, and a spread you must estimate.

A claim and twelve bottles

A bottling machine is claimed to fill bottles with a mean of 500 ml. Twelve bottles are taken from it and measured. Their mean is x̄ = 496 ml, and their standard deviation is s = 6 ml.

The hypotheses are about μ, the mean fill of every bottle the machine makes: H₀: μ = 500 and H₁: μ ≠ 500. The question is whether the machine is off in either direction, so the test is two-tailed, at the 5 percent level.

A z-test would need σ, the standard deviation of every bottle the machine makes. Nobody has given it. The only measure of spread is s, worked out from the same twelve bottles.

An estimated spread

s is the sample standard deviation with divisor n − 1: the squared deviations from x̄ are added and divided by 11, not by 12, and s is the square root of that.

Twelve bottles give only a rough estimate of σ. Some samples underestimate it, and then the gap from μ₀ looks larger than it is. So the statistic made with s varies more than z does, and it follows a t distribution instead of the standard normal.

A t distribution is centered on 0 and symmetric, like the standard normal, but lower in the middle and heavier in the tails. Which t distribution depends on the degrees of freedom.

t

The gold curve is the t distribution with 11 degrees of freedom. The dashed curve is the standard normal. The t curve is lower at 0 and stays higher out in both tails.

The statistic and its degrees of freedom

The statistic has the same shape as z, with s in place of σ: t = (x̄ − μ₀)/(s/√n). The top is the gap between the sample mean and the claimed mean. The bottom is the estimated standard error.

It has n − 1 degrees of freedom. The twelve deviations from x̄ always add up to zero, so once eleven of them are known, the twelfth is fixed. Working out s uses x̄, and that costs one degree of freedom: 12 bottles give 11.

The twelve bottles

The estimated standard error is s/√n = 6/√12 = 6/3.464 = 1.732 ml.

The gap is 496 − 500 = −4 ml, so t = −4/1.732 = −2.309, which the lesson rounds to −2.31. The sample mean is 2.31 estimated standard errors below the claimed mean.

t−2.312.20

The t distribution with 11 degrees of freedom. The shaded tail begins at −2.20 and holds 2.5 percent of the area; the same 2.5 percent lies above 2.20, at the other end. The bottles’ t = −2.31 lies just inside the shaded tail.

The verdict

With 11 degrees of freedom, P(T ≥ 2.201) = 0.025. For a two-tailed test at 5 percent, 2.5 percent goes in each tail, so the critical values are −2.201 and 2.201: reject H₀ if |t| > 2.201.

|t| = 2.31 > 2.20, so reject H₀. Stated about the machine: there is evidence at the 5 percent level that its mean fill is not 500 ml.

The p-value is 2 × P(T ≥ 2.309) = 0.0413, below 0.05. At the 1 percent level the critical value is 3.106, and 2.31 falls short of it, so the same twelve bottles would not be enough evidence at 1 percent.

One tail

If the only question had been whether the machine underfills, H₁: μ < 500, all 5 percent would go in the lower tail. With 11 degrees of freedom, P(T ≤ −1.796) = 0.05, so H₀ is rejected if t < −1.796.

t = −2.31 < −1.796, so H₀ is rejected here too, with p-value P(T ≤ −2.309) = 0.0207, half the two-tailed one. The tails are chosen from the question before the data are seen, never after.

t or z

When σ is given, the z-test applies and the two-tailed 5 percent critical value is 1.96. When the sample must estimate the spread with s, the t-test applies, on n − 1 degrees of freedom.

The smaller the sample, the more the two differ. The two-tailed 5 percent critical value of t is 2.201 for 11 degrees of freedom, 2.042 for 30 and 1.984 for 100. As the sample grows, s becomes a close estimate of σ and t gets closer to z.

Like the z-test, the t-test assumes the population is roughly normal. With twelve bottles there are too few for the central limit theorem to make up for a population that is far from normal.

The usual mistakes

Dividing by s instead of s/√n. −4/6 = −0.67 compares the gap with the spread of single bottles; the test is about the mean of twelve.

Using n degrees of freedom. Twelve bottles give 11, because s is worked out from x̄.

Using 1.96. That critical value belongs to the normal distribution; with 11 degrees of freedom it is 2.201.

Saying H₀ is proved when |t| falls short. Not rejecting means only that the sample did not give enough evidence against the claimed mean.

Nine batteries

In the application below, a consumer group tests 9 batteries against a claimed mean life of 20 hours. The variance is estimated from the nine lives with divisor 8, the test is one-tailed, and the statistic is compared with t on 8 degrees of freedom at 5 percent and at 1 percent.

Worked example: Nine Batteries Tested Against a Manufacturer's Claim of 20 Hours

Question A manufacturer claims that its batteries last 20 hours on average in a torch. A consumer group suspects they last less, and tests 9 batteries, which last 17.6, 18.0, 18.4, 18.6, 19.0, 19.4, 20.2, 20.4 and 21.2 hours. Battery lives are assumed to be normally distributed. (a) Find the sample mean, the unbiased estimate of the variance and the test statistic. (b) Test the claim at the 5% level, and say whether the conclusion would be the same at the 1% level.

  1. 1.Let μ be the mean life of the batteries. H0: μ = 20 and H1: μ < 20, a one-tailed test. The variance is unknown and the sample is small, so the test uses the t distribution with 9 − 1 = 8 degrees of freedom.

    17182122claim 20H0: mean 20, H1: mean < 20sd unknown, n = 9: use 8 df
    17182122claim 20H0: mean 20, H1: mean < 20sd unknown, n = 9: use 8 df
    The nine battery lives, in hours, against the claimed mean of 20.
  2. 2.The nine lives add up to 172.8, so x = 172.89 = 19.2 hours. Their deviations from 19.2 are −1.6, −1.2, −0.8, −0.6, −0.2, 0.2, 1.0, 1.2 and 2.0 hours, and their squares add up to 11.52.

    17182122claim 20mean 19.2mean = 172.8/9 = 19.2 hourssquared deviations add up to 11.52
    17182122claim 20mean 19.2mean = 172.8/9 = 19.2 hourssquared deviations add up to 11.52
    The sample mean is 19.2 hours, and the squared deviations from it add up to 11.52.
  3. 3.(a) s2 = 11.528 = 1.44, so s = 1.2 hours, and the test statistic is 19.2 − 201.2 / √9 = −0.80.4 = −2.0.

    0test statistic −2.08 df, if H0 is truevariance = 11.52/8 = 1.44, sd 1.2(19.2 − 20)/(1.2/3) = −2.0
    0test statistic −2.08 df, if H0 is truevariance = 11.52/8 = 1.44, sd 1.2(19.2 − 20)/(1.2/3) = −2.0
    (a) s2 = 11.528 = 1.44, and the test statistic is 19.2 − 201.2 / 3 = −2.0 on t(8).
  4. 4.(b) For a lower-tailed test at 5% on 8 degrees of freedom the critical value is −1.860. Since −2.0 < −1.860, reject H0: there is evidence at the 5% level that the batteries last less than 20 hours on average.

    −1.8600test statistic −2.08 df, if H0 is true−2.0 < −1.860: reject H0 at 5%
    −1.8600test statistic −2.08 df, if H0 is true−2.0 < −1.860: reject H0 at 5%
    (b) The lowest 5% of t(8) lies below −1.860, and −2.0 is in it: reject H0.
  5. 5.At the 1% level the critical value is −2.896, and −2.0 > −2.896, so H0 is not rejected: at 1% there is not enough evidence against the manufacturer's claim. Check: the deviations add up to 0, as deviations from a mean must.

    −2.8960test statistic −2.08 df, if H0 is true−2.0 > −2.896: do not reject at 1%
    −2.8960test statistic −2.08 df, if H0 is true−2.0 > −2.896: do not reject at 1%
    The lowest 1% lies below −2.896, and −2.0 is outside it: at 1%, H0 is not rejected.

Answer: (a) x = 19.2 hours, s2 = 1.44, and the test statistic is −2.0; (b) −2.0 < −1.860, so at 5% reject H0: there is evidence that the batteries last less than 20 hours on average; at 1% the critical value is −2.896 and H0 is not rejected

Common mistakes

  • Comparing −2.0 with the normal critical value −1.645. With σ estimated from only 9 batteries the statistic has the heavier-tailed t(8) distribution, and its critical value is −1.860.
  • Dividing the sum of squares by 9, which gives 1.28, or dividing s by 9 instead of by √9 = 3. The unbiased variance divides by n − 1 = 8, and the standard error is s√n = 0.4.

More statistical inference problems, worked step by step →

Practice The One-Sample T-Test in the app