Sampling · applications

Applications: Sampling

10 question types · Pre-University · each worked step by step with a figure that follows the steps

SAT · H2

01

A Survey of School Lunches Sampled Year Group by Year Group, with Shares That Must Be Whole Pupils

methodTake the Same Fraction of Every Stratum, Then Round by the Largest Remainders So That the Sample Still Adds Up

A school of 800 pupils has four year groups: 235 pupils in Year 7, 205 in Year 8, 190 in Year 9 and 170 in Year 10. The canteen wants a stratified sample of 50 pupils for a survey about school lunches, with each year group represented in proportion to its size. (a) Find the exact share of the sample for each year group, and show that rounding each share to the nearest whole number does not give a sample of 50. (b) Give every year group the whole-number part of its share, then give the pupils still needed to the year groups with the largest remainders. How many pupils are sampled from each year group?

schoolY7 235Y8 205Y9 190Y10 17080050 out of 800: 1 in every 16each share: the group size divided by 16
The school is split into its four year groups, the strata. The sampling fraction is 50800 = 116.
The sampling fraction is 50800 = 116: one pupil in every 16 is sampled, from every year group alike.
step 1 of 5

A stratified sample splits the population into groups, called strata, and samples each group in proportion to its size, so that every year group has the same weight in the sample as it has in the school. Each share is the size of the group multiplied by the sampling fraction 50800. The shares are rarely whole numbers, and rounding each one on its own can leave the total one too high or one too low.

  1. The sampling fraction is 50800 = 116: one pupil in every 16 is sampled, from every year group alike.
  2. Divide each year group by 16: Year 7 gets 23516 = 14.6875, Year 8 gets 20516 = 12.8125, Year 9 gets 19016 = 11.875 and Year 10 gets 17016 = 10.625. Check: 14.6875 + 12.8125 + 11.875 + 10.625 = 50.
  3. (a) Rounding each share to the nearest whole number gives 15 + 13 + 12 + 11 = 51 pupils, one more than the 50 the survey is for. Every share was rounded up, and the four small increases add up to more than a whole pupil.
  4. Give each year group the whole-number part of its share: 14 + 12 + 11 + 10 = 47, so 3 more pupils are needed. The remainders are 0.6875, 0.8125, 0.875 and 0.625. The three largest belong to Years 9, 8 and 7, and each of those year groups gets one more pupil.
  5. (b) The sample is 15 pupils from Year 7, 13 from Year 8, 12 from Year 9 and 10 from Year 10. Check: 15 + 13 + 12 + 10 = 50, and every year group is within one pupil of its exact share.

answer(a) the shares are 14.6875, 12.8125, 11.875 and 10.625, and rounding each one gives 51 pupils; (b) 15 from Year 7, 13 from Year 8, 12 from Year 9 and 10 from Year 10

techniqueSampling Methods · Populations and Samples

examsSAT · GCSE Higher · H2

Common pitfalls

  • Taking the same number from each year group, 12.5 each. That ignores the sizes of the groups: Year 7 has 65 more pupils than Year 10, so it needs a larger part of the sample for the sample to reflect the school.
  • Taking the extra pupil away from Year 7 because it is the largest year group. The fair place to take it from is the year group whose share was rounded up the most, Year 10, which gained 0.375 of a pupil by rounding.
02

A Town's Exercise Habits Asked Outside a Gym, and the Same Question Put to Every 120th Name on the Register

methodAsk Who Could Be Chosen and Whether They Are Like the Whole Town; a Systematic Sample Takes Every kth Name After a Random Start

A town council wants to know what fraction of the 24000 adults in the town exercise at least once a week. A volunteer asks 200 people outside a gym, and 176 of them say that they do. (a) Explain why this sample is biased, and say whether it makes the fraction look too large or too small. (b) The council instead takes a systematic sample of 200 from the electoral register of all 24000 adults, starting at the 37th name. Find the sampling interval and the positions of the first three names and the last name chosen. Given that 92 of these 200 adults exercise weekly, estimate the number of adults in the town who do.

gym176 exercise2488%only people at a gym can be asked176 of 200 say yes
Everyone asked was outside a gym, so adults who never exercise had no chance of being chosen.
Only people who are at a gym could be chosen for the first sample, and most people at a gym exercise. Adults who never exercise had no chance of being asked.
step 1 of 5

A sample is biased when the way it is chosen makes some members of the population more likely to be picked than others, in a way that is linked to the answer. A systematic sample lists the whole population, splits the list into equal blocks, chooses a random start in the first block, and then takes every kth name, where k is the size of the population divided by the size of the sample.

  1. Only people who are at a gym could be chosen for the first sample, and most people at a gym exercise. Adults who never exercise had no chance of being asked.
  2. (a) The sample is biased. It gives 176200 = 88%, which makes the fraction who exercise look too large. Asking more people at the same gym would not help, because the bias comes from where the sample is taken, not from its size.
  3. For the systematic sample, the interval is k = 24000200 = 120. The register is split into 200 blocks of 120 names, and one name is taken from each block.
  4. Start at the 37th name and add 120 each time. The first three names chosen are the 37th, 157th and 277th, and the last is the 37 + 199 × 120 = 23917th. Check: the last block runs from the 23881st name to the 24000th, and 23917 is in it.
  5. (b) The register sample gives 92200 = 0.46, so an estimated 0.46 × 24000 = 11040 adults exercise weekly, far fewer than the 88% from the gym suggests.

answer(a) it is biased, because only people at a gym can be asked, and the 88% it gives is too large; (b) the interval is 120; the 37th, 157th and 277th names, and the last is the 23917th; an estimated 11040 adults

techniqueSampling Methods · Populations and Samples

examsSAT · GCSE Higher · H2

Common pitfalls

  • Believing that a larger sample at the gym would remove the bias. 2000 people asked at the gym would give an answer just as far from the truth; a larger sample only reduces the chance variation, not the bias.
  • Dividing the wrong way round and taking an interval of 20024000. The interval is the number of names on the register for each name chosen: 24000 names shared among 200 choices is 120.
03

A Revision App and Exam Passes, First Compared by Who Chose to Use It and Then Tested by Random Allocation

methodAsk Who Decided Who Got the Treatment: When People Choose, a Confounding Variable Can Make the Gap; When Random Allocation Chooses, It Cannot

A school looks at last year's results. Of the 120 students who chose to use a revision app, 84 passed their mathematics exam; of the 180 students who did not use it, 99 passed. (a) Is this an observational study or an experiment? Find the pass rate of each group and the difference between them. (b) Name a confounding variable that could explain part of the difference. This year, 200 volunteers are split at random into two groups of 100: the group given the app has 66 passes and the other group has 60. Find the difference now, and say what it suggests.

last year: the students choseused app84 passed3670%no app99 passed8155%the school only recorded choicesan observational study
The students decided for themselves whether to use the app, so this is an observational study.
The students chose for themselves whether to use the app, and the school only recorded what happened. This is an observational study.
step 1 of 5

In an observational study the researchers record what people already do; in an experiment the researchers decide who receives the treatment. When people choose for themselves, anything that affects both the choice and the result is a confounding variable, and it can make a treatment look better than it is. Random allocation spreads such variables evenly between the groups, including the ones nobody thought to measure.

  1. The students chose for themselves whether to use the app, and the school only recorded what happened. This is an observational study.
  2. (a) The pass rates are 84120 = 70% for the students who used the app and 99180 = 55% for the others, a difference of 15 percentage points. An observational study shows that using the app and passing go together; it cannot show that the app caused the passes.
  3. A confounding variable affects both the choice and the result. Students who are more motivated are more likely to download a revision app, and they are also more likely to pass whether or not they use it. Motivation, or the time a student spends revising, is a confounding variable here.
  4. In the new study a random draw decides who gets the app, so motivated students are as likely to be in one group as in the other. The pass rates are 66100 = 66% with the app and 60100 = 60% without it.
  5. (b) The difference is now 66% − 60% = 6 percentage points. Most of last year's 15 points therefore came from which students chose the app, not from the app itself. What is left is small: on two groups of 100, a gap of 6 points could arise by chance alone, so this study shows at most a modest effect.

answer(a) an observational study; pass rates of 70% and 55%, a difference of 15 percentage points; (b) motivation, or the time spent revising, is a confounding variable; with random allocation the difference is 6 percentage points, so most of the first study's gap came from who chose the app, and a gap that size on 100 students each could be chance

techniqueObservational Studies and Experiments

examsSAT

Common pitfalls

  • Reading the 15-point gap as the effect of the app. The students who chose it were not like the students who did not, so the gap mixes the effect of the app with the effect of motivation.
  • Letting the students, or their teachers, pick the groups in the new study. Only random allocation makes the two groups alike in everything except the app.
04

Fish in a Lake Counted by Tagging and Recatching, and What Happens When Tags Fall Off

methodSet the Tagged Share of the Second Catch Equal to the Tagged Share of the Lake, Then Ask Which Way a Broken Assumption Moves the Estimate

To estimate the number of fish in a lake, an ecologist catches 120 fish, tags them and releases them. A week later she catches 150 fish, and 24 of them are tagged. (a) Estimate the number of fish in the lake. (b) She later learns that one tag in five falls off within a week. Say whether the estimate in (a) is too high or too low, and give a corrected estimate.

lake120 taggedN − 120Nsecond catch: 150 fish, 24 taggedlake: 120 tagged out of Ncatch: 24/150 = 0.16 tagged
The lake holds 120 tagged fish among N. In the second catch, 24 of the 150 fish, 16%, are tagged.
Let N be the number of fish in the lake. The share of the lake that is tagged is 120N, and the share of the second catch that is tagged is 24150 = 0.16.
step 1 of 5

The method assumes that the tagged fish mix freely with the others, and that the number of tagged fish does not change between the two catches. Then the share of tagged fish in the second catch estimates the share of tagged fish in the whole lake, and the one unknown in that equation is the number of fish.

  1. Let N be the number of fish in the lake. The share of the lake that is tagged is 120N, and the share of the second catch that is tagged is 24150 = 0.16.
  2. Set the two shares equal: 120N = 24150. Multiply both sides by 150N: 24N = 120 × 150 = 18000, so N = 1800024 = 750.
  3. (a) There are about 750 fish in the lake. Check: 120 tagged fish out of 750 is 120750 = 0.16, the same share as in the catch.
  4. If one tag in five falls off, only 45 × 120 = 96 fish still carry a tag at the second catch. Fewer tagged fish are caught than the lake size would give, and dividing by 120 rather than 96 makes the lake look larger than it is.
  5. (b) The estimate in (a) is too high. With 96 tagged fish, 96N = 24150 gives N = 96 × 15024 = 600 fish.

answer(a) about 750 fish; (b) too high: with 96 tags still in place, the estimate is 600 fish

techniqueCapture and Recapture · Populations and Samples

examsSAT · GCSE Higher

Common pitfalls

  • Adding the two catches, 120 + 150 = 270 fish. Some fish may have been caught twice, and the lake certainly holds fish that were never caught; the estimate comes from the shares, not from the totals.
  • Deciding that lost tags make the estimate too low. A fish that has lost its tag is counted as untagged, so the tagged share of the catch is too small, and a smaller share makes the estimate of the lake larger.
05

Cars Owned by the Five Households on a Short Street, and Every Sample of Two That Could Be Chosen

methodList Every Possible Sample, Work Out Each Sample Mean and Count How Often Each Mean Occurs; the Mean of the Sample Means Is the Population Mean

The five households on a short street own 0, 1, 1, 2 and 3 cars. A researcher chooses 2 of the households at random, without replacement, and works out X, the mean number of cars in her sample. (a) List all the possible samples and find the sampling distribution of X. Hence find P(X ≥ 2). (b) Find E(X), and compare it with the mean number of cars per household on the street.

A 0, B 1, C 1, D 2, E 3 carsAB 0.5AC 0.5AD 1AE 1.5BC 1BD 1.5BE 2CD 1.5CE 2DE 2.510 samples, all equally likely
There are 52 = 10 samples of two households, each with its own mean.
Call the households A, B, C, D and E, with 0, 1, 1, 2 and 3 cars. The ten samples and their means are AB 0.5, AC 0.5, AD 1, AE 1.5, BC 1, BD 1.5, BE 2, CD 1.5, CE 2 and DE 2.5.
step 1 of 5

The five households are the population, and its mean μ is a fixed number. A sample mean X changes from one sample to the next. Its sampling distribution lists every value that X can take, with its probability. Here it is found by listing every sample, since each of the 52 = 10 samples is equally likely.

  1. Call the households A, B, C, D and E, with 0, 1, 1, 2 and 3 cars. The ten samples and their means are AB 0.5, AC 0.5, AD 1, AE 1.5, BC 1, BD 1.5, BE 2, CD 1.5, CE 2 and DE 2.5.
  2. Count how often each mean occurs. X takes the values 0.5, 1, 1.5, 2 and 2.5 with probabilities 210, 210, 310, 210 and 110. Check: 2 + 2 + 3 + 2 + 1 = 10.
  3. (a) P(X ≥ 2) = 210 + 110 = 310, from the samples BE, CE and DE.
  4. Weight each value by its probability: E(X) = 0.5 × 2 + 1 × 2 + 1.5 × 3 + 2 × 2 + 2.5 × 110 = 1410 = 1.4.
  5. (b) The mean for the street is μ = 0 + 1 + 1 + 2 + 35 = 75 = 1.4 cars, so E(X) = μ. The sample mean is an unbiased estimator of the population mean, even though no single sample has a mean of exactly 1.4.

answer(a) X is 0.5, 1, 1.5, 2 or 2.5 with probabilities 210, 210, 310, 210 and 110, and P(X ≥ 2) = 310; (b) E(X) = 1.4, the same as the mean for the street, 1.4 cars per household

techniqueThe Sampling Distribution · Mean and Variance of X-bar · Populations and Samples

examsH2

Common pitfalls

  • Treating the two households with one car as one household and listing fewer samples. B and C are different households, so AB and AC are two different samples, and leaving one out gives the wrong probabilities.
  • Expecting a sample mean to equal the population mean. Each sample mean is 0.5, 1, 1.5, 2 or 2.5; it is the average over all the possible samples that equals 1.4.
06

Eggs from a Farm Weighed One at a Time and Twenty-Five at a Time

methodThe Sample Mean Has the Population Mean and the Population Variance Divided by n; Standardize with the Standard Deviation of the Mean, Not of One Egg

The masses of eggs from a farm are normally distributed with mean 62 g and standard deviation 5 g. An inspector weighs a random sample of 25 eggs and works out their mean mass, X grams. (a) Find E(X) and Var(X). (b) Using Φ(2) = 0.9772 and Φ(0.4) = 0.6554, find the probability that the mean mass of the sample is less than 60 g, and compare it with the probability that a single egg is less than 60 g.

6260gmean of 25one egg: sd 5mean of the sample: 62 g
The mean of 25 eggs is centered on the same 62 g as one egg: E(X) = 62.
The mean of the sample has the same expected value as a single egg: E(X) = μ = 62 g.
step 1 of 5

For a random sample of n from a population with mean μ and variance σ2, E(X) = μ and Var(X) = σ2n. When the population is normal, X is normal too, for any n. The mean of many eggs varies much less than a single egg does, because the light and heavy eggs in a sample balance each other.

  1. The mean of the sample has the same expected value as a single egg: E(X) = μ = 62 g.
  2. (a) Var(X) = σ2n = 5225 = 2525 = 1. So E(X) = 62 and Var(X) = 1, and the standard deviation of the mean is √1 = 1 g.
  3. The masses are normal, so X ∼ N(62, 1). Standardize with the standard deviation of the mean: z = 60 − 621 = −2, so P(X < 60) = 1 − Φ(2) = 1 − 0.9772 = 0.0228.
  4. For a single egg the standard deviation is 5 g: z = 60 − 625 = −0.4, so P(X < 60) = 1 − Φ(0.4) = 1 − 0.6554 = 0.3446.
  5. (b) The mean of the sample is below 60 g with probability 0.0228, while a single egg is below 60 g with probability 0.3446, about 15 times as likely. Check: the curve for the mean is 5 times as narrow, so 60 g is 2 of its standard deviations below 62 g, rather than 0.4 of them.

answer(a) E(X) = 62 g and Var(X) = 1; (b) P(X < 60) = 0.0228, against 0.3446 for a single egg

techniqueMean and Variance of X-bar · Standard Error

examsH2

Common pitfalls

  • Standardizing the sample mean with σ = 5 and getting 0.3446. That is the probability for a single egg; the mean of 25 eggs has standard deviation 5√25 = 1 g.
  • Dividing the standard deviation by 25 instead of by √25. It is the variance that is divided by n; the standard deviation is divided by √n = 5.
07

A Survey of Teenagers' Screen Time, Sized So That Both of Its Standard Errors Are Small Enough

methodWrite the Standard Error as a Function of n, Set It No Larger Than the Target, and Solve for the Smallest Whole Number n

A researcher is planning a survey of teenagers. Earlier studies suggest that daily screen time has a standard deviation of 48 minutes, and that about 40% of teenagers use a phone after midnight. (a) How many teenagers must she survey for the standard error of the mean screen time to be at most 4 minutes? (b) The same survey will estimate the proportion who use a phone after midnight, and that standard error must be at most 0.02. How many teenagers are needed for this, and how many must the survey include to meet both targets?

nSE04SE = 48/√nSE of the mean = 48/√nit must be at most 4
The standard error of the mean falls as n grows, but only as 1√n. The dashed line is the target of 4 minutes.
For the mean, the standard error is 48√n minutes, and the target is 48√n ≤ 4.
step 1 of 5

The standard error of an estimate is its standard deviation from one sample to the next. For a mean it is σ√n, and for a proportion it is √p(1 − p)n. Both get smaller as n grows, but only in step with √n: to halve a standard error takes four times the sample.

  1. For the mean, the standard error is 48√n minutes, and the target is 48√n ≤ 4.
  2. Multiply both sides by √n and divide both sides by 4: √n ≥ 12, so n ≥ 144.
  3. (a) She must survey at least 144 teenagers. Check: 48√144 = 4812 = 4 minutes exactly.
  4. For the proportion, the standard error is √0.4 × 0.6n = √0.24n. Square both sides of √0.24n ≤ 0.02: 0.24n ≤ 0.0004, so n ≥ 0.240.0004 = 600.
  5. (b) The proportion needs at least 600 teenagers, and a survey of 600 meets both targets. Check: √0.24600 = √0.0004 = 0.02, and with n = 600 the standard error of the mean is 48√600 ≈ 1.96 minutes, well under 4.

answer(a) at least 144 teenagers; (b) at least 600 for the proportion, so the survey must include 600 teenagers

techniqueStandard Error

examsSAT · H2

Common pitfalls

  • Solving 48n ≤ 4 and getting n ≥ 12. The standard error divides by √n, not by n, so a sample of 12 has a standard error of 48√12 ≈ 13.9 minutes.
  • Taking the smaller of the two sample sizes. A survey of 144 meets the first target but gives a standard error of √0.24144 ≈ 0.041 for the proportion, about twice the target; the survey must be large enough for both.
08

Skewed Service Times at a Post Office Counter, and the Mean Time for Sixty-Four Customers

methodHowever Skewed the Population, the Mean of a Large Sample Is Close to Normal: Use the Mean and the Standard Deviation over the Square Root of n, and Check That n Is Large Enough

The time a post office clerk takes to serve a customer has mean 4 minutes and standard deviation 4 minutes. The distribution is strongly skewed: most customers take a minute or two, and a few take far longer. (a) Using Φ(1.6) = 0.9452, find the probability that the mean service time of a random sample of 64 customers is more than 4.8 minutes. (b) A trainee suggests using the same method for the mean of just 4 customers. Using Φ(2) = 0.9772, find the probability that the method would then assign to a mean service time below 0 minutes, and explain what this shows.

04minutesone customer: mean 4, sd 4one customer: skewed, not normal64 customers: n is large
A single service time is strongly skewed: most are short and a few are very long.
The service times are not normal, but n = 64 is large, so by the central limit theorem X is approximately normal. Its mean is μ = 4 minutes and its standard deviation is σ√n = 4√64 = 48 = 0.5 minutes.
step 1 of 5

The central limit theorem says that for a large sample the mean X is approximately normal, X ∼ N(μ, σ2n) approximately, whatever the shape of the population. How large is large depends on the population: n ≥ 30 is a rule of thumb, and a strongly skewed population needs a larger sample than a symmetric one.

  1. The service times are not normal, but n = 64 is large, so by the central limit theorem X is approximately normal. Its mean is μ = 4 minutes and its standard deviation is σ√n = 4√64 = 48 = 0.5 minutes.
  2. Standardize: z = 4.8 − 40.5 = 1.6.
  3. (a) P(X > 4.8) = 1 − Φ(1.6) = 1 − 0.9452 = 0.0548.
  4. For 4 customers the standard deviation of the mean would be 4√4 = 2 minutes. A normal curve with mean 4 and standard deviation 2 puts 0 minutes at z = 0 − 42 = −2, so it gives P(X < 0) = 1 − Φ(2) = 1 − 0.9772 = 0.0228.
  5. (b) The method gives a probability of 0.0228 to a negative mean time, which is impossible. For a population this skewed, 4 customers are far too few for the mean to be close to normal. With 64 customers, 0 minutes is 8 standard deviations below the mean, and the normal curve puts almost no area there.

answer(a) P(X > 4.8) = 0.0548; (b) 0.0228 for a negative mean time, which is impossible, so 4 customers are too few for the central limit theorem to apply

techniqueThe Central Limit Theorem · Applying the Theorem · Standard Error

examsH2

Common pitfalls

  • Using the normal distribution for a single customer, P(X > 4.8) = 1 − Φ(0.2). One service time is strongly skewed, not normal; the theorem is about the mean of many customers.
  • Using σ = 4 rather than σ√n = 0.5 for the mean, which gives z = 0.2 and a probability of about 0.42. The mean of 64 customers has a standard deviation 8 times smaller than one customer's time.
09

A Courier's Van Loaded with Fifty Parcels, and the Chance That the Load Is Over Its Limit

methodThe Total of n Independent Values Has Mean n Times the Mean and Variance n Times the Variance; for Large n It Is Approximately Normal Whatever One Value's Distribution

The parcels a courier carries have masses with mean 11 kg and variance 8 kg2, and the distribution is not normal: it leans towards the lighter parcels. A van is loaded with 50 parcels chosen at random, and its safe load is 600 kg. (a) Using Φ(2.5) = 0.9938, find the probability that the total mass of the parcels is more than 600 kg. (b) Using Φ(1.96) = 0.975, find the total mass that the 50 parcels exceed with probability only 0.025.

kgz5500total of 50: mean 550mean of the total: 50 × 11 = 550 kg
The mean of the total is 50 × 11 = 550 kg.
Let T be the total mass of the 50 parcels. Its mean is E(T) = 50 × 11 = 550 kg.
step 1 of 5

The total T of n independent values has E(T) = nμ and Var(T) = nσ2. The total is n times the sample mean, so the central limit theorem applies to it as well: for large n, T is approximately normal, whatever the distribution of a single parcel's mass.

  1. Let T be the total mass of the 50 parcels. Its mean is E(T) = 50 × 11 = 550 kg.
  2. The masses are independent, so their variances add: Var(T) = 50 × 8 = 400, and the standard deviation of T is √400 = 20 kg. With n = 50, the central limit theorem gives T ∼ N(550, 202) approximately.
  3. Standardize the safe load: z = 600 − 55020 = 5020 = 2.5.
  4. (a) P(T > 600) = 1 − Φ(2.5) = 1 − 0.9938 = 0.0062, about 6 loads in every 1000.
  5. (b) The top 2.5% of totals lies above z = 1.96, since Φ(1.96) = 0.975. That total is 550 + 1.96 × 20 = 550 + 39.2 = 589.2 kg. Check: 589.2 − 55020 = 1.96.

answer(a) P(T > 600) = 0.0062; (b) 589.2 kg

techniqueApplying the Theorem · The Central Limit Theorem

examsH2

Common pitfalls

  • Multiplying the standard deviation by 50, giving 50 × √8 ≈ 141 kg. That treats the fifty parcels as copies of one parcel; for independent parcels it is the variance that is multiplied by 50, because light and heavy parcels partly balance each other.
  • Refusing to use the normal distribution because one parcel's mass is not normal. The question is about the total of 50 parcels, and the central limit theorem makes that total approximately normal.
10

A Poll of 100 Voters on a Referendum Too Close to Call, Approximated by a Normal Curve

methodReplace B(n, p) by a Normal Curve with Mean np and Variance np(1 - p), Widen Each Whole Number to Half a Unit Either Side, and Check Against the Exact Sum

In a referendum, exactly half of the voters support a new bridge. A polling company asks 100 voters chosen at random, and X of them support it. (a) Use a normal approximation with a continuity correction, and Φ(1.5) = 0.9332, to find P(X ≥ 58). Compare it with the exact binomial probability, which is 0.0666 to 4 decimal places. (b) Use the same approximation, with Φ(0.1) = 0.5398, to find the probability that exactly 50 of the 100 voters support the bridge, and explain why the continuity correction is essential here.

405060B(100, 0.5) and N(50, 25)mean = 100 × 0.5 = 50Var = 100 × 0.5 × 0.5 = 25, sd 5
Each bar of B(100, 0.5) has width 1, and the normal curve with the same mean 50 and variance 25 runs through their tops.
X ∼ B(100, 0.5) has mean np = 50 and variance np(1 − p) = 100 × 0.5 × 0.5 = 25. Both np and n(1 − p) are 50, well above 5, so X is approximately N(50, 52).
step 1 of 5

When n is large and p is not close to 0 or 1, X ∼ B(n, p) is close to a normal distribution with the same mean np and variance np(1 − p). The binomial takes only whole numbers, each drawn as a bar of width 1, while the normal curve is continuous. So each whole number k is replaced by the interval from k − 0.5 to k + 0.5: this is the continuity correction.

  1. X ∼ B(100, 0.5) has mean np = 50 and variance np(1 − p) = 100 × 0.5 × 0.5 = 25. Both np and n(1 − p) are 50, well above 5, so X is approximately N(50, 52).
  2. The event X ≥ 58 includes the whole bar for 58, which starts at 57.5. With the continuity correction, P(X ≥ 58) ≈ P(Y > 57.5), where Y ∼ N(50, 25), and z = 57.5 − 505 = 1.5.
  3. (a) P(X ≥ 58) ≈ 1 − Φ(1.5) = 1 − 0.9332 = 0.0668. The exact binomial sum is 0.0666, so the approximation is out by only 0.0002.
  4. Exactly 50 is the single bar from 49.5 to 50.5. Standardize both edges: z = 49.5 − 505 = −0.1 and z = 50.5 − 505 = 0.1.
  5. (b) P(X = 50) ≈ Φ(0.1) − Φ(−0.1) = 2 × 0.5398 − 1 = 0.0796, which agrees with the exact value 10050 × 0.5100 = 0.0796. Without the correction the interval would have no width, and a continuous curve gives a probability of 0 to any single value.

answer(a) P(X ≥ 58) ≈ 0.0668, against the exact 0.0666; (b) P(X = 50) ≈ 0.0796, the same as the exact value to 4 decimal places; without the correction a single value would have probability 0

techniqueApproximating a Binomial · Continuity Correction

Common pitfalls

  • Using 58 itself as the edge: z = 58 − 505 = 1.6 gives 0.0548, which is too small. The bar for 58 belongs to the event, and it starts at 57.5.
  • Dividing by the variance, 25, instead of the standard deviation: z = 7.525 = 0.3. The standard deviation is √25 = 5.
Mr. Chalk Read the guide