Which results count against the null
A test rejects when the test statistic lands in the most extreme part of its distribution under . Which part counts as extreme is set by the alternative hypothesis , before the data are seen.
A machine fills bags with mean 500 g. If the worry is underfilling, : , and only a sample mean well below 500 counts as evidence. If the worry is overfilling, : , and only a high one counts. These are one-tailed tests. If any change matters, : , and a result far out on either side counts. That is a two-tailed test.
One tail: all 5% at one end
At the 5% significance level, a one-tailed test for an increase puts the whole 5% in the upper tail. The critical value is the z with 5% of the area to its right, which is 95% to its left: z = 1.645, since P(Z > 1.645) = 0.0500. Reject if z > 1.645.
A test for a decrease uses the lower tail instead, and by symmetry rejects if z < −1.645.
A one-tailed test at 5%: the shaded upper tail beyond z = 1.645 holds the whole 5%.
Two tails: 2.5% at each end
A two-tailed test at 5% must split the 5% between the two tails, 2.5% in each. The upper critical value then has 97.5% of the area to its left: z = 1.96, since P(Z > 1.96) = 0.0250. Reject if z > 1.96 or z < −1.96, that is, if |z| > 1.96.
z = 1.8: p = 0.0719 ≥ 0.05, so a result this far out is not surprising enough under H₀, and H₀ is not rejected; the bar is 1.96
Two-tailed: find the smallest z that rejects H₀ at 5%
z = 1.8 in a two-tailed test: the area beyond 1.8 on both sides is p = 0.0719, more than 0.05, so is not rejected; the critical values are marked. Switch to one-tailed and only the upper tail counts: p = 0.0359, below 0.05. Drag z to find where each kind of test starts to reject.
The critical values
Each critical value comes from reading the normal table backward. A two-tailed test at any level uses the one-tailed value for half that level: two-tailed at 10% puts 5% in each tail, so its critical value is 1.645, the one-tailed value at 5%.
At every level the two-tailed value sits further out, so a result in a given direction must be more extreme to reject . That is the price of watching both directions at once.
Critical values of z at three significance levels. In each column the two-tailed value is further out than the one-tailed value.
One z, two verdicts
A sample gives z = 1.8. In a one-tailed test for an increase at 5%, 1.8 > 1.645, so is rejected; the p-value is P(Z > 1.8) = 0.0359. In a two-tailed test at 5%, 1.8 < 1.96, so is not rejected; the p-value counts both tails, 2 × 0.0359 = 0.0719.
This is why the kind of test is chosen before the data are seen. Waiting to see which way the data lean, and then running a one-tailed test in that direction, rejects whenever |z| > 1.645. If is true, that happens with probability 0.1000, twice the 5% the test claims.
The usual mistakes
Using 1.645 for a two-tailed test at 5%. 1.645 leaves 5% in one tail; a two-tailed test allows each tail only 2.5%, which needs 1.96.
Using 1.96 for a one-tailed test at 5%. 1.96 leaves only 2.5% in the tail, so the test would really be at the 2.5% level.
Taking 2.58 for the 5% level. 2.576 is the two-tailed critical value at 1%.
Comparing a negative z with +1.645 in a test for a decrease. The lower tail’s critical value is −1.645, and z = −2 rejects because −2 < −1.645.