How surprising is the result?
A critical region is drawn before the data arrives. A p-value is worked out after: it measures how surprising the actual result would be if the null hypothesis were true.
Suppose a test of whether a mean has increased gives z = 2.2. Shade everything at least as extreme as that, which for this test is every value of 2.2 or more. The shaded area is the p-value: .
So if were true, a result this far out or further would turn up about 14 times in 1,000.
The gold curve is the distribution of z if is true. The shaded tail starts at the observed value, 2.2, and its area, 0.0139, is the p-value.
Compare p with the level
The rule is to reject when p is below the significance level. Here 0.0139 < 0.05, so reject at the 5 percent level: data like this would be unlikely if were true.
This is the same verdict the critical region gives. The region for this test is , and . A z beyond 1.645 has a smaller tail than 0.05, and a z short of it has a larger one, so p < 0.05 exactly when z is in the critical region.
The p-value says more than the verdict alone. 0.0139 is below 0.02 as well, so is rejected at the 2 percent level too. It is above 0.01, so at the 1 percent level is not rejected.
Had z come to 1.2, p would be . That is above 0.05, so do not reject : a result like this would happen more than one time in ten by chance alone.
z = 2.2: p = 0.0139 < 0.05, so a result this far out would happen under H₀ less than 1 time in 20, and H₀ is rejected; the bar is 1.645
Two-tailed: find the smallest z that rejects H₀ at 5%
It opens on a one-tailed test at z = 2.2, where the shaded tail is p = 0.0139. Drag z back toward 1.645 and p grows to 0.05 there. Switch to two tails and the area beyond −2.2 is shaded as well.
Two tails
In a two-tailed test, a result counts as extreme in either direction. For z = 2.2 that means every value of 2.2 or more and every value of −2.2 or less. The normal curve is symmetric, so the two tails are equal, and p = 2 × 0.0139 = 0.0278.
With 9 heads in 10 tosses of a coin, the one-tailed p-value counts 9 or 10 heads: 11 in 1,024, about 1 percent. The two-tailed p-value also counts 1 or 0 heads, another 11 in 1,024, so , about 2 percent.
What a p-value is not
A p-value is worked out by assuming that is true. It is the probability of data at least this extreme, given . It cannot also be the probability that is true, given the data: the calculation took as true from the start.
So p = 0.03 means that results this extreme would happen about 3 times in 100 if were true. It does not mean there is a 3 percent chance that is true, or a 97 percent chance that the effect is real.
A p-value also says nothing about how large an effect is. A test of 10,000 people, with , whose mean is 100.5 against a claimed 100 has a standard error of , so and p = 0.0004. The evidence that the mean is above 100 is strong, yet the difference is only half a point.
Stating the conclusion
A conclusion is a sentence about the thing tested, not only about . A test of whether a new mug keeps coffee hot longer gives p = 0.01. That is under 0.05, so is rejected, and the conclusion is: there is evidence at the 5 percent level that the new mug keeps coffee hot longer.
It is evidence, not proof. A small p makes the data surprising if is true; it does not make the data impossible.
The usual mistakes
Reversing the rule. A small p rejects . A p above the level, such as 0.08 at 5 percent, keeps .
Reading p as the chance that is true. p is worked out assuming , so it is the chance of the data if holds.
Saying is proved when p is large. A large p means only that the data does not contradict .
Leaving out the context. "Reject " is true but incomplete; the conclusion names the mug, the coin or the mean that was tested.