The claim as a model
A die is claimed to be fair. Someone suspects that it lands on six too often, so it is rolled 60 times and the sixes are counted.
If the claim is true, each roll gives a six with probability . There is a fixed number of rolls, 60, and the rolls are independent. So the number of sixes X follows the binomial distribution , and on average there are sixes.
The hypotheses are about the probability p of a six. : , the die is fair. : , the die favors sixes. Both are written down, and the 5 percent level chosen, before the die is rolled.
Shade everything at least as extreme
The 60 rolls give 17 sixes. Under , 10 were expected. The alternative says sixes come too often, so the results at least as extreme as 17 are 17, 18, 19 and every count up to 60.
Their total probability is . From a calculator or cumulative binomial tables, , so .
The first terms show where that comes from: P(X = 17) = 0.0090, P(X = 18) = 0.0043, P(X = 19) = 0.0019, P(X = 20) = 0.0008, and the terms after that are smaller still.
The gold line joins the probabilities of 0, 1, 2 and more sixes in 60 rolls of a fair die, highest at 10. The shaded bars are 17 sixes or more; their heights add up to 0.0164.
The tail is the p-value
That tail probability is the p-value of the test: p = 0.0164. If the die really is fair, 60 rolls give 17 or more sixes only about 16 times in 1,000.
It is not the probability that the die is fair. It was worked out by assuming the die is fair, so it measures how surprising 17 sixes would be from a fair die.
The verdict, about the die
Compare p with the level: 0.0164 < 0.05, so reject . Stated about the die itself: there is evidence at the 5 percent level that the die lands on six more often than 1 time in 6.
It is evidence, not proof. A fair die does sometimes give 17 sixes in 60 rolls, so the die is not certainly loaded.
The level matters. At the 1 percent level, 0.0164 > 0.01, so would not be rejected.
The same test as a critical region
The test can also be set up with a critical region before rolling. Add tail probabilities from the top until the next one would pass 5 percent: is below 0.05, and is above it. So the critical region is , and the actual significance level is 0.0339.
17 is in the region, so the verdict is the same: reject . The two methods always agree, because a count is in the region exactly when its tail probability is at most the level.
The lower tail, and both tails
If the suspicion were that sixes come too rarely, would be , and the results at least as extreme would be counts at the bottom. and , so the critical region at 5 percent would be .
For a test of whether the die is unfair in either direction, a two-tailed test, each tail may hold at most 2.5 percent. At the bottom, is within 0.025. At the top, is within it, while is not. So the region is or , and 17 sixes still rejects .
The usual mistakes
Using the probability of exactly 17. P(X = 17) = 0.0090 leaves out every count beyond 17. The p-value adds them all: .
Leaving the observed count out. starts one count too late. The count that was seen belongs in the tail.
Swapping n and p. The model is : 60 rolls, each with probability . describes 6 trials with a 1 in 60 chance each.
Reading p as the chance the die is fair. 0.0164 is the chance of 17 or more sixes from a fair die, not the chance that this die is fair.