P-Values

The chance of data this extreme, if H0 held.

How surprising is the result?

A critical region is drawn before the data arrives. A p-value is worked out after: it measures how surprising the actual result would be if the null hypothesis H₀ were true.

Suppose a test of whether a mean has increased gives z = 2.2. Shade everything at least as extreme as that, which for this test is every value of 2.2 or more. The shaded area is the p-value: p = P(Z ≥ 2.2) = 1 − Φ(2.2) = 1 − 0.9861 = 0.0139.

So if H₀ were true, a result this far out or further would turn up about 14 times in 1,000.

z2.2

The gold curve is the distribution of z if H₀ is true. The shaded tail starts at the observed value, 2.2, and its area, 0.0139, is the p-value.

Compare p with the level

The rule is to reject H₀ when p is below the significance level. Here 0.0139 < 0.05, so reject H₀ at the 5 percent level: data like this would be unlikely if H₀ were true.

This is the same verdict the critical region gives. The region for this test is z ≥ 1.645, and P(Z ≥ 1.645) = 0.05. A z beyond 1.645 has a smaller tail than 0.05, and a z short of it has a larger one, so p < 0.05 exactly when z is in the critical region.

The p-value says more than the verdict alone. 0.0139 is below 0.02 as well, so H₀ is rejected at the 2 percent level too. It is above 0.01, so at the 1 percent level H₀ is not rejected.

Had z come to 1.2, p would be P(Z ≥ 1.2) = 0.1151. That is above 0.05, so do not reject H₀: a result like this would happen more than one time in ten by chance alone.

−3−2−101231.645z = 2.2p = 0.0139p < 0.05: reject H₀two-tailedone-tailed

z = 2.2: p = 0.0139 < 0.05, so a result this far out would happen under H₀ less than 1 time in 20, and H₀ is rejected; the bar is 1.645

Two-tailed: find the smallest z that rejects H₀ at 5%

It opens on a one-tailed test at z = 2.2, where the shaded tail is p = 0.0139. Drag z back toward 1.645 and p grows to 0.05 there. Switch to two tails and the area beyond −2.2 is shaded as well.

Two tails

In a two-tailed test, a result counts as extreme in either direction. For z = 2.2 that means every value of 2.2 or more and every value of −2.2 or less. The normal curve is symmetric, so the two tails are equal, and p = 2 × 0.0139 = 0.0278.

With 9 heads in 10 tosses of a coin, the one-tailed p-value counts 9 or 10 heads: 11 in 1,024, about 1 percent. The two-tailed p-value also counts 1 or 0 heads, another 11 in 1,024, so p = 22/1024, about 2 percent.

What a p-value is not

A p-value is worked out by assuming that H₀ is true. It is the probability of data at least this extreme, given H₀. It cannot also be the probability that H₀ is true, given the data: the calculation took H₀ as true from the start.

So p = 0.03 means that results this extreme would happen about 3 times in 100 if H₀ were true. It does not mean there is a 3 percent chance that H₀ is true, or a 97 percent chance that the effect is real.

A p-value also says nothing about how large an effect is. A test of 10,000 people, with σ = 15, whose mean is 100.5 against a claimed 100 has a standard error of 15 ÷ √10000 = 0.15, so z = 0.5 / 0.15 = 3.33 and p = 0.0004. The evidence that the mean is above 100 is strong, yet the difference is only half a point.

Stating the conclusion

A conclusion is a sentence about the thing tested, not only about H₀. A test of whether a new mug keeps coffee hot longer gives p = 0.01. That is under 0.05, so H₀ is rejected, and the conclusion is: there is evidence at the 5 percent level that the new mug keeps coffee hot longer.

It is evidence, not proof. A small p makes the data surprising if H₀ is true; it does not make the data impossible.

The usual mistakes

Reversing the rule. A small p rejects H₀. A p above the level, such as 0.08 at 5 percent, keeps H₀.

Reading p as the chance that H₀ is true. p is worked out assuming H₀, so it is the chance of the data if H₀ holds.

Saying H₀ is proved when p is large. A large p means only that the data does not contradict H₀.

Leaving out the context. "Reject H₀" is true but incomplete; the conclusion names the mug, the coin or the mean that was tested.

Practice P-Values in the app