One number and its error
A machine fills bags of flour, and the standard deviation of a bag’s mass is known to be 10 g. To estimate the mean mass of all its bags, 100 bags are weighed, and their mean is x̄ = 502 g.
502 g is a point estimate of : a single number. Another 100 bags would give a slightly different mean, and the point estimate on its own says nothing about how far from it might be.
The standard error measures how far sample means typically fall from . For samples of 100 it is .
Plus or minus 1.96 standard errors
Sample means are normally distributed about , with standard deviation equal to the standard error. In the standard normal distribution, , so 2.5 percent lies beyond 1.96 in each tail and P(−1.96 < Z < 1.96) = 0.95.
So in 95 percent of samples, x̄ lands within 1.96 standard errors of . When it does, is within 1.96 standard errors of x̄. That gives the 95 percent confidence interval: standard error.
For the flour, , which runs from 502 − 1.96 = 500.04 g to 502 + 1.96 = 503.96 g.
The gold curve is the standard normal distribution. The shaded middle, from −1.96 to 1.96, holds 95 percent of the area, leaving 2.5 percent in each tail.
What the 95 percent means
is a fixed number, and the interval from 500.04 g to 503.96 g either contains it or does not. The 95 percent belongs to the method: if many samples of 100 bags were taken and an interval made from each, about 95 in every 100 of those intervals would contain .
The interval is about the mean of all the bags. It does not say that 95 percent of bags weigh between 500.04 g and 503.96 g. Single bags vary with a standard deviation of 10 g, far more than that.
Other levels of confidence
A different level uses a different z. For 90 percent, 5 percent is left in each tail, and , so the interval is , from 500.355 g to 503.645 g.
For 99 percent, 0.5 percent is left in each tail, and , so the interval is , from 499.424 g to 504.576 g.
More confidence needs a wider interval. The 99 percent interval is 2.576 ÷ 1.645, about 1.57 times as wide as the 90 percent one.
A larger sample, a narrower interval
The standard error divides by . With 400 bags instead of 100, it is , and if the mean were again 502 g the 95 percent interval would be , from 501.02 g to 502.98 g.
Four times as many bags halves the width, because . To halve the width again would take 1,600 bags.
Intervals and tests
Suppose the bags are labeled 500 g. That value lies below the 95 percent interval, 500.04 g to 503.96 g, so a mean of 500 g is not a plausible value at this level.
A two-tailed test of : at the 5 percent level reaches the same verdict: , which is beyond 1.96, so is rejected. A claimed mean outside the 95 percent interval is always rejected by this test, and one inside it never is.
The usual mistakes
Going one standard error either side. contains only about 68 times in 100. A 95 percent interval needs 1.96 standard errors.
Doubling the margin. The already reaches both ways, so the half-width is 1.96 × 1 = 1.96, not 3.92.
Using instead of the standard error. is a range for single bags, not for the mean.
Saying there is a 95 percent chance that is in this interval. The 95 percent describes how often the method succeeds, not this one interval.
Journey times on a bus route
In the application below, 64 journeys are timed and is estimated from the sample as 8 minutes, so the standard error is minute. The timetable’s 32 minutes is then set against the interval.
Worked example: Journey Times on a Bus Route, a 95% Confidence Interval for the Mean, and the Timetable's 32 Minutes
Question A bus company times 64 journeys on one route, chosen at random over a month. The mean time is 34.5 minutes, and the unbiased estimate of the standard deviation is 8 minutes. (a) Find a 95% confidence interval for the mean journey time on the route. (b) The timetable allows 32 minutes. Use the interval to comment on the timetable, and say what the interval does not tell you.
1.The sample is large, so by the central limit theorem the sample mean is approximately normal, and s = 8 can be used in place of σ. The standard error is 8√64 = 1 minute.
The sample mean is 34.5 minutes, and its standard error is 8√64 = 1 minute. 2.A 95% interval reaches 1.96 standard errors either side of the sample mean: 34.5 ± 1.96 × 1 = 34.5 ± 1.96.
A 95% interval reaches 1.96 standard errors either side of the sample mean. 3.(a) The 95% confidence interval is (32.54, 36.46) minutes.
(a) The 95% confidence interval is (32.54, 36.46) minutes. 4.32 minutes lies below the interval, so a mean of 32 minutes is not a plausible value at this level: the timetable allows too little time. A two-tailed test of μ = 32 at the 5% level would reject it.
The timetable's 32 minutes lies below the interval, so it is not a plausible value for the mean. 5.(b) The interval is about the mean journey time. It does not say that 95% of journeys take between 32.54 and 36.46 minutes: with a standard deviation of 8 minutes, single journeys vary far more than that. Nor is there a 95% chance that this one interval contains μ; the 95% describes how often intervals made this way contain the true mean.
(b) The interval is for the mean journey time, not for single journeys, and the 95% describes the method, not this one interval.
Answer: (a) (32.54, 36.46) minutes; (b) 32 is below the interval, so the timetable allows too little time; the interval is for the mean, not for single journeys, and the 95% describes how often intervals made this way contain the true mean
Common mistakes
- Going 1.96 × 8 either side of the mean, which gives (18.82, 50.18). That range is for single journeys; the interval for the mean uses the standard error 8√64 = 1.
- Saying that 95% of journeys take between 32.54 and 36.46 minutes. The interval estimates the mean of all journeys; a single journey is often more than 8 minutes from the mean.