Two claims about the population
A hypothesis test weighs two claims about a population parameter, such as a proportion p or a mean . The null hypothesis, , says there is no effect: the coin is fair, the mean is what it always was. The alternative hypothesis, , says there is an effect.
Both are written in symbols, about the parameter. For a coin suspected of landing heads too often, : p = 0.5 and : p > 0.5. For a machine that fills 500 g bags, suspected of underfilling, : and : . For a die suspected of showing six too often, : and : .
The claim the test is looking for evidence of, the bias or the underfilling, always goes in . The null is the position held until the data argue against it.
Why the null has one value
names a single value of the parameter, and that is what makes it testable. If p = 0.5, the number of heads in 10 tosses is X ~ B(10, 0.5), one exact distribution, so the chance of any result can be worked out.
: p > 0.5 names no single value. It covers p = 0.51 and p = 0.9 alike, and gives no one distribution to work with. So a test assumes and asks how surprising the data would be if it were true.
If : p = 0.5 is true, each of the 1,024 sequences of 10 tosses is equally likely. The bars count the sequences giving 0 to 10 heads: 252 give 5, and only 10 + 1 = 11 give 9 or more.
Evidence against the null
The coin lands heads 9 times in 10. If it is fair, the chance of 9 or more heads is , about 1%. A result that rare under is evidence against it.
The usual rule is to fix a significance level before seeing the data, often 5%, and reject if the probability of a result at least this extreme is below it. 0.0107 is below 0.05, so reject : there is evidence at the 5% level that the coin lands heads more often than tails.
Failing to reject is not proof
Suppose instead the coin lands heads 6 times. If it is fair, the chance of 6 or more heads is . That is not surprising, so is not rejected.
That does not show the coin is fair. If heads had probability 0.6, the chance of 6 or more heads would be 0.633, so 6 heads fits that coin even better. The data cannot tell the two apart; they are simply not surprising enough to reject .
So the conclusion is worded with care: there is not enough evidence that the coin is biased. A test rejects or fails to reject it, and never proves it true.
The usual mistakes
Putting the claim to be shown in . A drug company testing whether its drug works better writes : no difference and : the drug works better, because the effect needs the evidence.
Writing hypotheses about the sample. x̄ = 503 is a fact about the data, not a hypothesis. Hypotheses are about the population: : .
Saying that is accepted or proved. Not rejecting it means only that the data were not surprising enough under it.
Answering neither. Every statement in a test takes a side: “the mean is unchanged” is , and “the mean has increased” is .