Sampling Methods

Chosen to avoid bias, not for convenience.

What a sampling method is for

A sample is used to estimate something about a population, so it should be like the population. A method of choosing that tends to favor some members over others is biased, and a biased sample gives estimates that are wrong in the same direction every time.

The methods below differ in how the members are chosen. The question to ask of each one is whether every member had a fair chance of being chosen.

A simple random sample

In a simple random sample, every member of the population has the same probability of being chosen, and each choice is left to chance, independent of the others.

To choose 50 of a school’s 1,000 students, number the students from 1 to 1,000. Use a random number generator to produce numbers in that range, skipping any number that has already come up, until 50 different students are chosen. Each student has a probability of 50 ÷ 1,000 = 1/20 of being in the sample. Drawing 50 names from a hat holding all 1,000 does the same job.

Nobody decides who is in the sample, so the method favors no group. A simple random sample needs a full list of the population, called a sampling frame, and that list is not always available.

0102030405060708090100

A simple random sample of 10 from a list of 100, chosen by a random number generator: 5, 22, 24, 44, 46, 52, 60, 64, 71 and 84. The gaps are uneven, and two chosen names can be next to each other.

A stratified sample

Sometimes the population splits into groups that may answer differently, such as age groups or year groups. These groups are called strata. In a stratified sample, you split the population into its strata, then take a simple random sample from each stratum in proportion to its size.

A sports club has 240 members: 120 juniors, 80 adults and 40 seniors. A sample of 30 is 30 ÷ 240 = 1/8 of the club, so take 1/8 of each group: 120 ÷ 8 = 15 juniors, 80 ÷ 8 = 10 adults and 40 ÷ 8 = 5 seniors. Check that the strata add up: 15 + 10 + 5 = 30.

Every group is then represented in the same proportion as in the club, which a simple random sample only does on average. When the shares are not whole numbers, they have to be rounded so that the sample still adds up to its total; the first application below shows how.

juniorsadultsseniors

The stratified sample of 30 from the club: 15 juniors, 10 adults and 5 seniors, each 1/8 of its group of 120, 80 or 40.

A systematic sample

In a systematic sample, you take members at a fixed interval along a list. To choose 50 names from a list of 1,000, the interval is 1,000 ÷ 50 = 20. Choose a random starting point among the first 20 names, say the 7th, then take every 20th name after it: the 7th, 27th, 47th, and so on. The last name chosen is the 7 + 49 × 20 = 987th.

A systematic sample is quick to take, and it spreads the sample evenly along the list. But it relies on the order of the list being fair. Suppose a list of apartments goes floor by floor, with 10 apartments on each floor, and every 10th apartment on the list is a large corner apartment. An interval of 10 then picks the same position on every floor, and the sample could be all corner apartments, or none of them.

0102030405060708090100

A systematic sample of 10 from a list of 100: the interval is 100 ÷ 10 = 10, and the random start is the 4th name. Every gap is exactly 10.

A convenience sample

A convenience sample takes whoever is easy to reach. To find out how often the people of a town exercise, a volunteer stands outside a gym and asks the people going in and out.

Only people at the gym can be chosen, and most people at a gym exercise. Everyone who never exercises has no chance of being asked, so the sample overstates how much the town exercises. The venue has decided who could be sampled.

Asking more people at the same gym does not help. A larger sample reduces the chance variation from one sample to the next, but the bias comes from where the sample is taken, and it stays.

at the gymeveryone elsepeopleonly this part can be asked

The town’s adults, with the small part who are at the gym highlighted. A sample taken outside the gym can only come from that part.

A census

If you ask every member of the population, it is not a sample at all: it is a census. A census leaves nobody out, so it has no sampling error, and the numbers it gives are the parameters themselves.

A census is often impractical. A country’s census of every household takes years to plan and costs a great deal. And when the measuring destroys what is measured, as when matches are tested by striking them, a census would leave nothing to sell.

A quota sample

In a quota sample, the interviewer is given a fixed number to fill from each group, and chooses freely who to ask. An interviewer on a shopping street might be told to find 6 people under 30, 6 aged 30 to 60 and 6 over 60.

A quota sample is like a stratified sample without the randomness. Each group is filled, but within each group the interviewer picks whoever is willing and nearby. Anyone at work during the afternoon has no chance of being asked, and the interviewer may pass over people who look busy or unfriendly. Quota sampling is cheap and needs no list of the population, but it can be biased in ways that are hard to see.

The usual mistakes

Taking the same number from every stratum. A stratified sample takes from each group in proportion to its size, so the club’s 120 juniors get three times as many places as its 40 seniors.

Believing a larger sample removes bias. A sample of 2,000 people outside a gym is as biased as a sample of 200.

Dividing the wrong way for a systematic sample. The interval is the population size divided by the sample size, 1,000 ÷ 50 = 20, not 50 ÷ 1,000.

Calling a quota sample stratified. Both fill each group, but only the stratified sample chooses the members of each group at random.

Worked example: A Survey of School Lunches Sampled Year Group by Year Group, with Shares That Must Be Whole Pupils

Question A school of 800 pupils has four year groups: 235 pupils in Year 7, 205 in Year 8, 190 in Year 9 and 170 in Year 10. The canteen wants a stratified sample of 50 pupils for a survey about school lunches, with each year group represented in proportion to its size. (a) Find the exact share of the sample for each year group, and show that rounding each share to the nearest whole number does not give a sample of 50. (b) Give every year group the whole-number part of its share, then give the pupils still needed to the year groups with the largest remainders. How many pupils are sampled from each year group?

  1. 1.The sampling fraction is 50800 = 116: one pupil in every 16 is sampled, from every year group alike.

    schoolY7 235Y8 205Y9 190Y10 17080050 out of 800: 1 in every 16each share: the group size divided by 16
    schoolY7 235Y8 205Y9 190Y10 17050 out of 800: 1 in every 16each share: the group size divided by 16
    The school is split into its four year groups, the strata. The sampling fraction is 50800 = 116.
  2. 2.Divide each year group by 16: Year 7 gets 23516 = 14.6875, Year 8 gets 20516 = 12.8125, Year 9 gets 19016 = 11.875 and Year 10 gets 17016 = 10.625. Check: 14.6875 + 12.8125 + 11.875 + 10.625 = 50.

    schoolY7 235Y8 205Y9 190Y10 170800sample14.687512.812511.87510.6255014.6875 + 12.8125 + 11.875 + 10.625= 50
    schoolY7 235Y8 205Y9 190Y10 170sample14.687512.812511.87510.6255014.6875 + 12.8125 + 11.875 + 10.625= 50
    Each year group's exact share is its size divided by 16, and the four shares add up to 50.
  3. 3.(a) Rounding each share to the nearest whole number gives 15 + 13 + 12 + 11 = 51 pupils, one more than the 50 the survey is for. Every share was rounded up, and the four small increases add up to more than a whole pupil.

    schoolY7 235Y8 205Y9 190Y10 170800sample1513121151rounded: 15 + 13 + 12 + 11 = 51one pupil too many
    schoolY7 235Y8 205Y9 190Y10 170sample1513121151rounded: 15 + 13 + 12 + 11 = 51one pupil too many
    (a) Rounding every share to the nearest whole number gives 15 + 13 + 12 + 11 = 51, one pupil too many.
  4. 4.Give each year group the whole-number part of its share: 14 + 12 + 11 + 10 = 47, so 3 more pupils are needed. The remainders are 0.6875, 0.8125, 0.875 and 0.625. The three largest belong to Years 9, 8 and 7, and each of those year groups gets one more pupil.

    schoolY7 235Y8 205Y9 190Y10 170800sample14 + 112 + 111 + 11047 + 3whole parts: 14 + 12 + 11 + 10 = 473 more, to the 3 largest remainders
    schoolY7 235Y8 205Y9 190Y10 170sample14 + 112 + 111 + 11047 + 3whole parts: 14 + 12 + 11 + 10 = 473 more, to the 3 largest remainders
    The whole-number parts make 47. The remainders 0.875, 0.8125 and 0.6875 of Years 9, 8 and 7 are the largest, so those three each get one more pupil.
  5. 5.(b) The sample is 15 pupils from Year 7, 13 from Year 8, 12 from Year 9 and 10 from Year 10. Check: 15 + 13 + 12 + 10 = 50, and every year group is within one pupil of its exact share.

    schoolY7 235Y8 205Y9 190Y10 170800sample151312105015 + 13 + 12 + 10 = 50Year 10 has 10, not 11
    schoolY7 235Y8 205Y9 190Y10 170sample151312105015 + 13 + 12 + 10 = 50Year 10 has 10, not 11
    (b) The sample is 15, 13, 12 and 10 pupils, which adds up to 50.

Answer: (a) the shares are 14.6875, 12.8125, 11.875 and 10.625, and rounding each one gives 51 pupils; (b) 15 from Year 7, 13 from Year 8, 12 from Year 9 and 10 from Year 10

Common mistakes

  • Taking the same number from each year group, 12.5 each. That ignores the sizes of the groups: Year 7 has 65 more pupils than Year 10, so it needs a larger part of the sample for the sample to reflect the school.
  • Taking the extra pupil away from Year 7 because it is the largest year group. The fair place to take it from is the year group whose share was rounded up the most, Year 10, which gained 0.375 of a pupil by rounding.

More sampling problems, worked step by step →

Worked example: A Town's Exercise Habits Asked Outside a Gym, and the Same Question Put to Every 120th Name on the Register

Question A town council wants to know what fraction of the 24000 adults in the town exercise at least once a week. A volunteer asks 200 people outside a gym, and 176 of them say that they do. (a) Explain why this sample is biased, and say whether it makes the fraction look too large or too small. (b) The council instead takes a systematic sample of 200 from the electoral register of all 24000 adults, starting at the 37th name. Find the sampling interval and the positions of the first three names and the last name chosen. Given that 92 of these 200 adults exercise weekly, estimate the number of adults in the town who do.

  1. 1.Only people who are at a gym could be chosen for the first sample, and most people at a gym exercise. Adults who never exercise had no chance of being asked.

    gym176 exercise2488%only people at a gym can be asked176 of 200 say yes
    gym176 yes2488%only people at a gym can be asked176 of 200 say yes
    Everyone asked was outside a gym, so adults who never exercise had no chance of being chosen.
  2. 2.(a) The sample is biased. It gives 176200 = 88%, which makes the fraction who exercise look too large. Asking more people at the same gym would not help, because the bias comes from where the sample is taken, not from its size.

    gym176 exercise2488%176/200 = 88%biased: the fraction looks too large
    gym176 yes2488%176/200 = 88%biased: the fraction looks too large
    (a) The sample is biased: 88% makes the fraction who exercise look too large.
  3. 3.For the systematic sample, the interval is k = 24000200 = 120. The register is split into 200 blocks of 120 names, and one name is taken from each block.

    the register: 24000 names in 200 blocks of 1200120240360480120 names24000 names, 200 to chooseinterval: 24000 divided by 200 = 120
    register: 200 blocks of 1200120240360480120 names24000 names, 200 to chooseinterval: 24000 divided by 200 = 120
    The register is split into 200 blocks of k = 24000200 = 120 names, and one name is taken from each block.
  4. 4.Start at the 37th name and add 120 each time. The first three names chosen are the 37th, 157th and 277th, and the last is the 37 + 199 × 120 = 23917th. Check: the last block runs from the 23881st name to the 24000th, and 23917 is in it.

    the register: 24000 names in 200 blocks of 120012024036048037157277397start at 37, then add 120 each timelast: 37 + 199 × 120 = 23917
    register: 200 blocks of 120012024036048037157277397start at 37, then add 120 each timelast: 37 + 199 × 120 = 23917
    From a random start at 37, every 120th name: 37, 157, 277, and so on, to the 23917th.
  5. 5.(b) The register sample gives 92200 = 0.46, so an estimated 0.46 × 24000 = 11040 adults exercise weekly, far fewer than the 88% from the gym suggests.

    gym176 exercise2488%register92 exercise10846%92/200 = 0.460.46 × 24000 = 11040 adults
    gym176 yes2488%register92 yes10846%92/200 = 0.460.46 × 24000 = 11040 adults
    (b) The register sample gives 46%, so an estimated 11040 adults exercise weekly.

Answer: (a) it is biased, because only people at a gym can be asked, and the 88% it gives is too large; (b) the interval is 120; the 37th, 157th and 277th names, and the last is the 23917th; an estimated 11040 adults

Common mistakes

  • Believing that a larger sample at the gym would remove the bias. 2000 people asked at the gym would give an answer just as far from the truth; a larger sample only reduces the chance variation, not the bias.
  • Dividing the wrong way round and taking an interval of 20024000. The interval is the number of names on the register for each name chosen: 24000 names shared among 200 choices is 120.

More sampling problems, worked step by step →

Practice Sampling Methods in the app