A rule fixed before the data
A hypothesis test starts from the null hypothesis, , which usually says that nothing has changed. If is true, the test statistic z follows the standard normal distribution: values near 0 are common, and values far out in either direction are rare.
Before any data is collected, choose a significance level, often 5 percent. Then mark off the most extreme 5 percent of the distribution, on the side the alternative hypothesis points to. Those values of z are the critical region, and the value where the region begins is the critical value.
For a test of whether a mean has increased, the region is the top 5 percent. , so the critical value is 1.645 and the critical region is .
The gold curve is the standard normal distribution, the distribution of z if is true. The shaded tail from 1.645 holds 5 percent of the area: that tail is the critical region.
Where the result lands
Now collect the data and work out z. Suppose it comes to 2.5. That lies beyond 1.645, inside the critical region. If were true, a z of 2.5 or more would happen with probability , about 6 times in 1,000. A result that rare counts as evidence against , so reject at the 5 percent level.
Suppose instead that z comes to 1.2. That is short of 1.645, outside the region. Values of 1.2 or more happen with probability when is true, more than one time in ten, so do not reject .
Not rejecting does not prove it true. It means only that this data is not surprising enough to rule it out at the level chosen.
The same critical region, with two possible results marked. z = 2.5 lies under the shaded tail, so it rejects . z = 1.2 lies to the left of the shading, so it does not.
Which tail, and how far out
The alternative hypothesis decides where the region goes, and it must be chosen before the data is seen. For a test of whether a mean has decreased, the region is the lowest 5 percent: .
For a test of whether a mean has changed in either direction, a two-tailed test, the 5 percent is split into 2.5 percent in each tail. , so the region is or . The critical values sit further out than 1.645, because each tail holds only half as much.
A smaller significance level moves the critical values further out still. At 1 percent, a one-tailed test has critical value 2.326, since , and a two-tailed test has critical values , since 0.005 lies beyond each of them.
z = 1.5: p = 0.1336 ≥ 0.05, so a result this far out is not surprising enough under H₀, and H₀ is not rejected; the bar is 1.96
Two-tailed: find the smallest z that rejects H₀ at 5%
It opens on a two-tailed test with z at 1.5: the two shaded tails hold about 13 percent, too much to reject . Drag z out to 1.96, where they hold 5 percent, 2.5 percent in each; switch to one tail and the dashed critical value moves in to 1.645.
The region in the units of the data
A critical region can be written for the sample mean itself instead of z. A machine fills boxes of cereal with a mean of 500 g and a standard deviation of 12 g. A manager suspects that it has started to overfill, so : and : , at the 5 percent level, using a sample of 36 boxes.
The standard error is . The critical region is , and , so the region is .
The 36 boxes have a mean of 503.8 g. That is in the critical region, so reject : there is evidence at the 5 percent level that the machine is overfilling. As a check, , which is past 1.645.
When the result is a count
For a count, the critical region is a set of whole numbers, and its probability usually cannot be exactly 5 percent. A coin is tossed 10 times to test whether it favors heads, with : p = 0.5 and : p > 0.5. Under the number of heads X follows B(10, 0.5).
Add the probabilities from the top. , and . The first is below 0.05 and the second is above it, so the critical region is .
The probability of the region, 0.0107, is the actual significance level of the test: the chance of rejecting when the coin is in fact fair. It is below the 5 percent asked for, because the next value down, 8, would take the region over 5 percent.
The usual mistakes
Choosing the region after seeing the data. Once the result is known, a region can always be drawn around it. The level, the tail and the critical value are all fixed first.
Using the wrong critical value. 1.645 belongs to a one-tailed test at 5 percent and 1.96 to a two-tailed test at 5 percent. A two-tailed test with 1.645 puts 5 percent in each tail, 10 percent in all.
Letting a count’s region go over the level. For the coin, has probability 0.0547, which is more than 0.05, so 8 is not in the region, however close 0.0547 is to 5 percent.
Saying that is proved. A result outside the region fails to reject ; it does not show that is true.
A basketball player’s free throws
In the application below, a player who scored 40 percent of her free throws last season takes 20 throws after a summer of coaching. Under her score follows B(20, 0.4), and the critical region is found from the top, the same way as for the coin.
Worked example: A Basketball Player's Free Throws After a Summer of Coaching, Tested Against Last Season's Rate
Question Last season a basketball player scored 40% of her free throws. After a summer of coaching she believes she has improved, and she takes 20 free throws to find out. (a) Using a binomial model, find the critical region for a test at the 5% level of H0: p = 0.4 against H1: p > 0.4, and the actual significance level of the test. (b) She scores 12 of the 20. State the conclusion of the test in context.
1.Let X be the number of free throws she scores out of 20. Under H0, X ∼ B(20, 0.4). The alternative is p > 0.4, so the test is one-tailed and the critical region is at the top: X ≥ k for the smallest k with P(X ≥ k) ≤ 0.05.
Under H0 the number she scores is X ∼ B(20, 0.4). The alternative p > 0.4 puts the critical region in the top tail. 2.Add the binomial terms from the top: P(X ≥ 13) = 0.0210 and P(X ≥ 12) = 0.0565. The first is below 0.05 and the second is above it.
From the top, the tail from 13 has probability 0.0210, and the tail from 12 has 0.0565, which is more than 5%. 3.(a) The critical region is X ≥ 13. The actual significance level is P(X ≥ 13) = 0.0210, about 2.1%: that is the chance of rejecting H0 when in fact she has not improved.
(a) The critical region is the gold bars, X ≥ 13, and the actual significance level is 0.0210. 4.She scored 12, which is not in the critical region. In the same way, P(X ≥ 12) = 0.0565 is more than 0.05.
She scored 12, the bar just outside the critical region. 5.(b) Do not reject H0. There is not enough evidence at the 5% level that her success rate has risen above 40%. This does not show that she has not improved: 12 out of 20 is 60%, but 20 throws are too few to tell that rate apart from chance.
(b) Do not reject H0: there is not enough evidence that her rate has risen above 40%.
Answer: (a) the critical region is X ≥ 13, and the actual significance level is 0.0210; (b) 12 is not in the critical region, since P(X ≥ 12) = 0.0565, so do not reject H0: there is not enough evidence that her success rate has risen above 40%
Common mistakes
- Using P(X = 13) = 0.0146, the chance of exactly 13, in place of P(X ≥ 13). The test asks how likely a result this extreme or more extreme is, so every term from 13 to 20 is added.
- Taking X ≥ 12 as the critical region because 0.0565 is close to 5%. The region's probability must not exceed the level, so 12 is left out and the region starts at 13.