Sampling

Stage 16 of 23 Strand 3 of 5 11 lessons

11 illustrated lessons, each teaching the why before the how.

Revise Sampling with flashcards →

Jump to a lesson

Populations and Samples

A parameter is fixed; a statistic varies.

A parameter describes the whole population and a statistic describes only your sample

There are a thousand people in the population, and you can only ask thirty.

The population mean is fixed but unknown. You can compute the sample mean.

So a statistic is an estimate of a parameter, and it changes with each sample.

Now you

Is the true proportion of faulty parts a parameter or a statistic?

Is the proportion faulty in the box you opened a parameter or a statistic?

Sampling Methods

Chosen to avoid bias, not for convenience.

A sampling method is chosen to avoid bias rather than to be convenient

In a simple random sample, every member has the same probability of being chosen, and the selections are independent.

In a stratified sample, split the population into groups, then take from each group in proportion to its size.

In a systematic sample, take every tenth name on the list. It is fast, but it relies on the order of the list being fair.

A convenience sample takes whoever is easy to reach. At a gym, the venue decides who can be sampled.

Ask every member instead and it is a census — complete, and often impractical.

In a quota sample, fill a fixed count from each group, choosing freely. It is stratified sampling without the randomness.

Now you

A survey about exercise asks the first 20 people leaving a gym. What is wrong?

A school asks every single student about lunch. What is that called?

Observational Studies and Experiments

Assigning the treatment is what buys a cause.

Only a study that assigns the treatment at random can support a claim about cause

In an experiment the researcher assigns the treatment — here by a coin toss.

In an observational study you only record groups people put themselves into.

Self-chosen groups can differ before the treatment, and that alone can explain a gap.

Assigning at random spreads every other difference evenly over the two groups.

The control group is treated identically except for the treatment being tested.

Random assignment supports a cause; observing without it supports only an association.

Random sampling settles who a result covers; random assignment settles whether the result is a cause.

Now you

Volunteers were split by coin toss. What does that give the study?

An observational study finds a new drug goes with lower blood pressure. What may you conclude?

Capture and Recapture

The tagged share of a catch mirrors the pond.

Matching the tagged fraction of a catch to the pond estimates the whole population

Tag 20 fish and release them — they scatter through a pond of unknown size.

Recapture 30 fish: 6 carry tags — a fifth of the catch is tagged.

If the catch is representative of the pond, the fractions match — solving gives about 100 fish.

The estimate rests on assumptions: tagged fish mix fully, and every tag stays on.

Now you

Suppose the tagged fish spread right through the pond before the next catch. Is the estimate sound?

Suppose half the tags fall off before the next catch. Is the estimate sound?

The Sampling Distribution

Every possible sample mean, piled up.

Take many samples and their means form a distribution of their own

The population is ten scores with a mean of 5. One sample of four gives a mean of 6.

Take two more samples of four from the same ten scores, and each one has a different mean.

Keep only the means and they pile up between 4 and 6, where the scores ran 2 to 8.

Now you

Samples of 4 are taken from these ten scores again and again. How are the sample means spread?

Samples of 3 are taken from these ten scores again and again. How are the sample means spread?

Mean and Variance of X-bar

Unbiased, with variance divided by n.

The sample mean is right on average and its variance is divided by the sample size

On average the sample mean equals the population mean, so it is unbiased.

Independent variances add to nσ², and dividing by n divides variance by .

So every extra observation narrows the estimate a little further.

Now you

The population variance is 4 and the sample size is 9. What is the variance of the sample mean?

The population variance is 25 and the sample size is 4. What is the variance of the sample mean?

Standard Error

Halve it and you need four times the data.

The standard error is the standard deviation of the sample mean, so it falls with the root of n

Take the square root of the variance and the n comes out as a root: σ/√n.

The standard error falls fast at small n and then very slowly as n grows larger.

To halve the standard error you need four times as many readings.

Now you

Sigma is 12 and n is 4. What is the standard error?

Sigma is 6 and n is 25. What is the standard error?

The Central Limit Theorem

Means go normal whatever the population was.

Sample means become normal as the sample grows, whatever shape the population had

This population is lopsided and uneven, nothing like a bell.

Take means of samples from it and they pile up symmetrically anyway.

That is the central limit theorem, and it is why the normal curve appears so often.

Now you

As the sample size grows, what does the central limit theorem say about the population itself?

As the sample size grows, what does the central limit theorem say about the sample means?

Applying the Theorem

Thirty is a rule of thumb, not a sharp edge.

A sample of about thirty is usually enough for the normal approximation to be usable

Once the sample means are close enough to normal, every method for the normal distribution applies to them.

Thirty is a rule of thumb, not a sharp cutoff.

A badly lopsided population needs more than thirty, not exactly thirty.

Now you

Is a sample of 5 usually enough for the normal approximation?

Is a sample of 30 usually enough for the normal approximation?

Approximating a Binomial

Match the mean and variance to a normal.

A binomial with enough trials is close enough to a normal to be treated as one

Binomial bars already look bell-shaped once n is reasonably large.

Give the curve the same mean np and variance np(1 − p) as the bars, and it sits on top of them.

The fit needs room in both tails: np and n(1 − p) must both be greater than 5.

Now you

X ~ B(150, 1/5). Which normal approximates it?

X ~ B(75, 1/5). Which normal approximates it?

Continuity Correction

Bars have width; a curve does not.

Swapping bars for a curve needs half a unit of width added back at each edge

A binomial bar for X = 7 really covers everything from 6.5 to 7.5.

The curve has no bars, so you must say where the bar edges were.

P(X ≥ 7) keeps the whole bar for 7, so the boundary is its left edge, 6.5.

P(X > 7) drops that bar, so the boundary moves out to its right edge, 7.5.

Now you

Approximating P(X < 7), which boundary do you use?

Approximating P(X ≥ 5), which boundary do you use?

Continue your journey in the app — save your progress