Type One and Type Two Errors

Crying wolf, or missing a real effect.

Two ways to be wrong

A test ends in one of two verdicts, reject H₀ or do not reject it, and H₀ is in fact either true or false. That gives four outcomes, and two of them are errors.

Rejecting H₀ when it is true is a type one error: a false alarm. Not rejecting H₀ when it is false is a type two error: a real effect missed. The other two outcomes, rejecting a false H₀ and keeping a true one, are correct decisions.

Take a machine that fills cereal boxes with a mean of 500 g and a standard deviation of 12 g. A sample of 144 boxes is weighed to test H₀: μ = 500 against H₁: μ > 500, at the 5 percent level. A type one error stops a machine that was filling correctly. A type two error leaves a machine running that has started to overfill.

The chance of a type one error

The standard error is 12 ÷ √144 = 12 / 12 = 1 g, so z = (x̄ − 500) / 1 = x̄ − 500. The critical region is z ≥ 1.645, which is x̄ ≥ 501.645 g.

A type one error can only happen when H₀ is true. Then z follows the standard normal distribution, and the chance it lands in the critical region is P(Z ≥ 1.645) = 0.05. So the probability of a type one error is the significance level itself, chosen before the test. It is written α.

z1.645

The gold curve is the distribution of z if H₀ is true. The shaded area beyond 1.645 is the chance of rejecting H₀ when it is true, the type one error, α = 0.05.

The chance of a type two error

A type two error can only happen when H₀ is false, so its probability depends on what the true mean is. Suppose the machine has drifted to a mean of 503 g. Then x̄ is normal with mean 503 and standard error 1, so z = x̄ − 500 is normal with mean 3 and standard deviation 1.

The test fails to reject when z < 1.645, the same line as before. The probability of that, written β, is the area under the true distribution to the left of the line: β = P(z < 1.645) = Φ(1.645 − 3) = Φ(−1.355) = 0.0877.

So about 9 times in 100, a sample from this overfilling machine would pass the test. The chance of detecting the drift, 1 − β = 0.9123, is called the power of the test.

z1.645

The gold curve is the distribution of z under H₀, centered at 0. The dashed curve is its distribution when the true mean is 503 g, centered at 3. The shaded area under the dashed curve, to the left of 1.645, is β = 0.0877: results from the true distribution that fall short of the line.

Demanding more evidence

To make false alarms rarer, lower the significance level. At 1 percent the critical value moves out to 2.326, so α = 0.01. The line has moved toward the true distribution, and more of it now falls short: β = Φ(2.326 − 3) = Φ(−0.674) = 0.2502.

Move the line further, to 2.5, and α = P(Z ≥ 2.5) = 0.0062 while β = Φ(2.5 − 3) = Φ(−0.5) = 0.3085. With the sample size fixed, every move of the line that shrinks one error grows the other.

z2.5

The same two curves with the line moved out to 2.5. The tail of the gold curve beyond the line is now only 0.0062, but the shaded area under the dashed curve has grown to β = 0.3085.

A larger sample lowers both

The sample size changes how far apart the two curves sit. With 36 boxes instead of 144, the standard error is 12 ÷ √36 = 2 g, so a true mean of 503 g puts the center of z at (503 − 500) ÷ 2 = 1.5 instead of 3.

At the 5 percent level α is still 0.05, but β = Φ(1.645 − 1.5) = Φ(0.145) = 0.5576: more than half the time, the test would miss the drift. With 144 boxes it was 0.0877. A larger sample separates the two curves, which is the one way to reduce β without raising α.

The usual mistakes

Swapping the two names. Type one is rejecting a true H₀, the false alarm. Type two is keeping a false H₀, the miss.

Taking β as 1 − α. 1 − 0.05 = 0.95 is the chance of keeping H₀ when H₀ is true, a correct decision. β is worked out under the true mean, with H₀ false.

Shading under the wrong curve. β is an area under the true distribution, the one centered at 3, not under the H₀ curve.

Reading α as the chance that H₀ is true after a rejection. α is the chance of rejecting, given that H₀ is true.

Loaves that are light

In the application below, an inspector weighs 25 loaves to test whether their mean is below 400 g. The critical region is found under H₀, and the chance of a type two error is then found under the true mean of 395 g, as for the cereal.

Worked example: A Bakery's 400 g Loaves Inspected by Weighing 25, and the Chance That Light Loaves Pass the Test

Question A bakery's loaves are labeled 400 g, and their masses are normally distributed with standard deviation 10 g. An inspector weighs 25 loaves and will test H0: μ = 400 against H1: μ < 400 at the 5% level. (a) Find the critical region for the sample mean, and state the probability of a Type I error. (b) In fact the mean mass is 395 g. Find the probability that the test fails to detect this, and say what kind of error that is.

  1. 1.Under H0 the mean of 25 loaves is normal with mean 400 g and standard deviation 10√25 = 2 g. The alternative is μ < 400, so the test is one-tailed with the critical region in the lower tail.

    400if H0 is true: mean 400sd of the mean = 10/√25 = 2 gH1: mean < 400, the lower tail
    400if H0 is true: mean 400sd of the mean = 10/√25 = 2 gH1: mean < 400, the lower tail
    If H0 is true, the mean of 25 loaves is normal with mean 400 g and standard deviation 2 g.
  2. 2.The lowest 5% of the standard normal distribution lies below z = −1.645, so the critical value is 400 − 1.645 × 2 = 396.71 g.

    396.71400if H0 is true: mean 400400 − 1.645 × 2 = 396.71
    396.71400if H0 is true: mean 400400 − 1.645 × 2 = 396.71
    The lowest 5% of that curve lies below 400 − 1.645 × 2 = 396.71 g.
  3. 3.(a) The critical region is x < 396.71 g. A Type I error is rejecting H0 when it is true, and its probability is the significance level, 0.05.

    396.71400if H0 is true: mean 400reject H0 if the mean < 396.71 gP(Type I error) = 0.05
    396.71400if H0 is true: mean 400reject H0 if the mean < 396.71 gP(Type I error) = 0.05
    (a) The shaded tail is the critical region x < 396.71; its area, 0.05, is the probability of a Type I error.
  4. 4.With the true mean 395 g, the sample mean is normal with mean 395 and standard deviation 2. The test fails to reject H0 when x ≥ 396.71, which is z = 396.71 − 3952 = 0.855 standard deviations above 395.

    396.71400if H0 is true: mean 400395in fact: mean 395true mean 395: missed if the mean >= 396.71z = (396.71 − 395)/2 = 0.855
    396.71400if H0 is true: mean 400395in fact: mean 395true mean 395: missed if the mean >= 396.71z = (396.71 − 395)/2 = 0.855
    In fact the mean is 395 g. The test misses this whenever x lands to the right of the same line, 396.71 g.
  5. 5.(b) P(x ≥ 396.71) = 1 − 0.8037 = 0.196. This is a Type II error, failing to reject H0 when it is false: about one inspection in five would pass loaves that are on average 5 g light. Weighing more loaves would make this probability smaller.

    396.71400if H0 is true: mean 400395in fact: mean 395P(Type II error) = 1 − 0.8037= 0.196
    396.71400if H0 is true: mean 400395in fact: mean 395P(Type II error) = 1 − 0.8037= 0.196
    (b) The shaded area under the lower curve is P(x ≥ 396.71) = 0.196, the probability of a Type II error.

Answer: (a) the critical region is x < 396.71 g, and the probability of a Type I error is 0.05; (b) 0.196, the probability of a Type II error

Common mistakes

  • Using the standard deviation of one loaf, 10 g, to find the critical value, which gives 400 − 16.45 = 383.55 g. The test uses the mean of 25 loaves, whose standard deviation is 2 g.
  • Answering that the probability of a Type II error is 1 − 0.05 = 0.95. That is the probability of not rejecting H0 when H0 is true; a Type II error happens when H0 is false, so it is worked out with the true mean of 395 g.

More statistical inference problems, worked step by step →

Practice Type One and Type Two Errors in the app