Standardizing a sample mean
For one value from , counts the standard deviations between the value and the mean. A test about a mean uses the sample mean x̄ instead, and the spread of x̄ is not but the standard error . So the test statistic is : the number of standard errors between the sample mean and the mean claimed by .
Sugar is sold in bags labeled 500 g. The filling machine’s standard deviation is known from long use to be 12 g. An inspector weighs 36 bags and finds a mean of 496 g. The standard error is , so . The sample mean is 2 standard errors below the label.
The whole test
Hypotheses: is the machine’s mean fill. The question is whether it has drifted either way, so : and : , a two-tailed test at 5%.
Test statistic: under , x̄ ~ , so .
Compare: the two-tailed critical values at 5% are , and −2 < −1.96, so z is in the critical region. Equally, the p-value is 2 × P(Z < −2) = 2 × 0.0228 = 0.0455, which is below 0.05.
Conclusion, in context: reject . There is evidence at the 5% level that the machine’s mean fill is no longer 500 g.
If the inspector’s question had been only whether the bags are underfilled, : , the critical value would be −1.645 and the p-value P(Z < −2) = 0.0228. The conclusion is the same.
z = −1.5: p = 0.1336 ≥ 0.05, so a result this far out is not surprising enough under H₀, and H₀ is not rejected; the bar is 1.96
Two-tailed: find the smallest z that rejects H₀ at 5%
A two-tailed test at z = −1.5: p = 0.1336, so stands. Drag z to the sugar bags’ −2: the area beyond 2 standard errors on both sides is p = 0.0455, below 0.05, and z lies beyond the critical value −1.96.
When the z-test applies
The test needs , the population standard deviation, to be known, as it is for a machine with a long record. It also needs x̄ to be normal: either the population is normal, or n is large enough for the central limit theorem to make x̄ approximately normal anyway.
A large sample sees small differences
Suppose the machine’s mean fill really is 499 g, 1 g under the label, and the sample happens to show exactly that. With 36 bags, the standard error is 2 g and , with p-value 2 × P(Z < −0.5) = 0.6171. Nothing is detected.
With 900 bags, the standard error is , and the same 1 g gives , with p-value 0.0124. Now is rejected at 5%.
The difference is the same; only the standard error shrank. So a significant result says the difference is unlikely to be chance, not that it is large: a huge sample can detect a difference too small to matter.
Under , the mean of 36 bags follows the wide dashed curve, , and the mean of 900 bags the narrow gold curve, . A mean of 499 is well inside the wide one, but 2.5 standard errors out on the narrow one.
The usual mistakes
Leaving the gap undivided. 496 − 500 = −4 is the gap in grams; z counts standard errors, so it is .
Dividing by instead of the standard error. is the z-score of one bag; the test is about the mean of 36 bags.
Dividing by n. is not the standard error; the variance is divided by n, so is divided by .
Using a one-tailed p-value in a two-tailed test. With : , a mean as far above 500 would count just as much, so P(Z < −2) = 0.0228 is doubled to 0.0455.
Worked example: A Coffee Machine Checked After a Repair, and a Two-Tailed Test of Whether Its Mean Pour Has Changed
Question A coffee machine is set to pour 250 ml into each cup, and the volume it pours is normally distributed with standard deviation 6 ml. After a repair, a sample of 36 cups has a mean volume of 252.1 ml. (a) Test at the 5% level whether the mean volume has changed: state the hypotheses, and find the test statistic and the p-value. (b) Would the conclusion be the same at the 1% level? Give both conclusions in context.
1.Let μ be the mean volume after the repair. The question is whether it has changed in either direction, so the test is two-tailed: H0: μ = 250 and H1: μ ≠ 250.
If H0 is true, the standardized sample mean follows N(0, 1). A change in either direction counts, so both tails matter. 2.Under H0 the mean of 36 cups is normal with mean 250 and standard deviation 6√36 = 1 ml. The test statistic is z = 252.1 − 2501 = 2.1.
The mean of 36 cups has standard deviation 6√36 = 1 ml, so z = 2.1. 3.The p-value counts results at least this far from 250 in either direction: 2 × P(Z > 2.1) = 2 × (1 − 0.98214) = 0.0357.
The p-value is the shaded area in both tails beyond ± 2.1: 0.0357. 4.(a) 0.0357 < 0.05, or equally 2.1 > 1.96, so reject H0. There is evidence at the 5% level that the mean volume has changed from 250 ml since the repair.
(a) The 5% critical region is beyond ± 1.96, and z = 2.1 lies in it: reject H0. 5.(b) At the 1% level the critical values are ± 2.576, and 0.0357 > 0.01, so do not reject H0. At 1% there is not enough evidence that the mean volume has changed. The same data can be significant at one level and not at another, which is why the level is chosen before the data are seen.
(b) The 1% critical region is beyond ± 2.576, and z = 2.1 is outside it: do not reject H0.
Answer: (a) H0: μ = 250, H1: μ ≠ 250; z = 2.1 and the p-value is 0.0357, below 0.05 (2.1 > 1.96): reject H0, there is evidence that the mean volume has changed; (b) no: 0.0357 > 0.01 (2.1 < 2.576), so at the 1% level there is not enough evidence of a change
Common mistakes
- Standardizing with the standard deviation of one cup: 252.1 − 2506 = 0.35. The test is about the mean of 36 cups, whose standard deviation is 6√36 = 1 ml.
- Reporting P(Z > 2.1) = 0.0179 as the p-value. The alternative is μ ≠ 250, so a result as far below 250 would count just as much, and the tail is doubled.