Point Estimates

Unbiased means right on average, not right.

One number for an unknown value

A number that describes a whole population, like its mean μ or the proportion p of voters who support a plan, is a parameter. It is fixed, but usually unknown, because the whole population cannot be measured.

A sample gives a statistic, a number worked out from the sample. Used as a single best guess at a parameter, it is a point estimate. The sample mean x̄ is the point estimate of μ, and the sample proportion p̂ is the point estimate of p.

Five bags from a filling machine weigh 498, 503, 501, 497 and 506 g. Their total is 2505 g, so x̄ = 2505/5 = 501 g, and 501 g is the estimate of the machine’s mean fill. In a poll, 18 of 60 voters support a plan, so p̂ = 18/60 = 0.3 estimates the proportion of all voters.

The rule, “add up the sample and divide by n”, is the estimator. The number it gives for one particular sample, 501 g, is the estimate.

Right on average

A different sample gives a different estimate, so an estimator has a distribution of its own. It is unbiased if the mean of that distribution is the parameter: over all possible samples, its estimates average out to the true value.

Check it on a population small enough to list every sample. The population is the three values 1, 3 and 5, so μ = 3. Draw a sample of two, with replacement. There are 3 × 3 = 9 equally likely samples: (1, 1), (1, 3), (1, 5), (3, 1), (3, 3), (3, 5), (5, 1), (5, 3) and (5, 5).

Their means are 1, 2, 3, 2, 3, 4, 3, 4 and 5. Only three of the nine equal 3, but the others miss by the same amounts above and below, and the average of all nine is (1 + 2 + 3 + 2 + 3 + 4 + 3 + 4 + 5)/9 = 27/9 = 3, exactly μ. So x̄ is unbiased.

12345

The means of all nine samples of two from 1, 3 and 5, one dot each. They balance at 3, the population mean.

An estimator that leans

Not every sensible rule is unbiased. To estimate the largest value in the population, 5, the obvious rule is the largest value in the sample. For the nine samples it gives 1, 3, 5, 3, 3, 5, 5, 5 and 5.

It can never be more than 5 and is often less, so on average it falls short: (1 + 3 + 5 + 3 + 3 + 5 + 5 + 5 + 5)/9 = 35/9 = 3.89. This estimator is biased: it underestimates on average.

135

The largest value of each of the nine samples, one dot each. None is above 5, so they balance below it, at 3.89.

How far off?

Unbiased does not mean right. Six of the nine sample means of 1, 3 and 5 miss μ, and in a real sample the estimate is almost never exactly the parameter. A point estimate on its own says nothing about how large its miss is likely to be.

That is measured by the spread of the estimator. For x̄ it is the standard error σ/√n. If 100 bags have mean mass 502 g and σ = 10 g, the standard error is 10/√100 = 1 g, so 502 g is likely to be within about 2 g of μ. A range built this way is a confidence interval.

The usual mistakes

Reading unbiased as correct. An unbiased estimator is right on average over many samples, and any one of its estimates can still miss.

Expecting an estimator that is always exactly right. Chance decides which values land in the sample, so every estimate carries some error.

Mixing up the parameter and the estimate. μ is a fixed property of the population; x̄ is worked out from a sample and changes from sample to sample.

Practice Point Estimates in the app