Observational Studies and Experiments

Assigning the treatment is what buys a cause.

Two ways to study a treatment

Does a new drug help patients recover? Does a revision app raise exam marks? Questions like these compare a group that had a treatment with a group that did not. What matters is who decided which group each person was in.

An experiment

In an experiment, the researcher assigns the treatment. Sixty volunteers agree to take part, and for each one a coin is tossed: heads, the new drug; tails, a dummy pill that looks the same but contains no drug.

The researcher decided who got the drug, and decided it by chance. That is what makes this an experiment.

60 volunteerscoinnew drugcoindummy pill

In the experiment, a coin toss sends each volunteer to the new drug or to the dummy pill.

An observational study

In an observational study, nobody is assigned anything. The researcher finds patients who already take the drug and patients who do not, and records how each group does. The groups were formed by the patients’ own choices, or their doctors’, before the study began.

100 patientsown choicealready take itown choicealready do not

In the observational study, the patients are already in their groups, by their own choice.

Self-chosen groups can differ from the start

Groups that people choose for themselves can differ before any treatment. Suppose 100 patients are studied. Of the 40 who chose to take the drug, 30 were in good health to begin with and 10 were frail. Of the 60 who did not, 24 were in good health and 36 were frail. Healthier patients may be more likely to look after themselves, which includes taking a new drug.

Now suppose the drug does nothing at all, and that 5 out of every 6 healthy patients recover while half of the frail patients do. In the drug group, 25 of the 30 healthy patients and 5 of the 10 frail ones recover: 30 out of 40, which is 75%. In the other group, 20 of the 24 healthy patients and 18 of the 36 frail ones recover: 38 out of 60, about 63%.

The drug group recovers more often, by about 12 percentage points, and the drug did none of it. The gap comes entirely from the health the patients started with. A variable like this, which affects both who takes the treatment and how they do, is called a confounding variable.

take itdo nothealthier firstfrailer firstthe groups already differ

The 100 patients, with the 40 who chose the drug on the left. Three quarters of them, 30, were healthier from the start, against 24 of the 60, or 40%, in the other group.

Random assignment evens out the rest

In an experiment, a coin toss decides each patient’s group, and the toss does not depend on anyone’s health. Had the same 100 patients been split by coin tosses, about half of the 54 healthy ones would land in each group. Chance might make it 25 and 29 rather than 27 and 27, but nothing would push the healthy patients toward one group.

The same is true of every other difference between patients, including the ones nobody thought to measure: age, diet, how carefully they follow instructions. Random assignment spreads them all roughly evenly over the two groups, and the larger the groups, the more evenly. Then the treatment is the one thing that differs between the groups in a planned way.

drugdummyhealthier firstfrailer firstchance splits every other difference evenly

After a coin toss, the drug group and the dummy group hold about the same share of healthier patients, and of frailer ones.

The control group

The group given the treatment is the treatment group. The group it is compared with is the control group. The control group is treated in exactly the same way in every respect but one: the treatment being tested.

That is why the control group is given a dummy pill, called a placebo, rather than nothing. Patients who believe they are being treated can feel better for that reason alone, so both groups take a pill, see the same doctors and have the same checkups. Often neither the patients nor the doctors know who has the real drug until the end, so that nobody’s expectations can treat the two groups differently.

Cause or association

When the treatment was assigned at random and the treatment group does better, a cause is supported: chance sorted the groups, so nothing but the treatment, and chance itself, can explain the gap.

When the groups were only observed, a gap shows an association: the treatment and the result go together in the data. It cannot show that the treatment caused the result, because something else, such as the patients’ health to begin with, may have sorted the groups.

Who the result covers

Two different random choices answer two different questions. Random sampling, choosing who takes part at random from a population, settles who the result covers: it can be extended to the population the sample came from. Random assignment, deciding at random who gets the treatment, settles whether the result is a cause.

The 60 volunteers in the experiment were not a random sample of anyone; they chose to take part. So the experiment can show that the drug helps these volunteers. Whether it helps patients in general is a separate question, about how well the volunteers stand for them.

The usual mistakes

Reading a gap in an observational study as the effect of the treatment. The 75% against about 63% above comes from who chose the drug.

Letting the participants, or the researchers, pick the groups in an experiment. Only random assignment makes the groups alike in everything except the treatment.

Thinking random assignment makes the result cover everyone. It supports a cause for the people in the study; covering a whole population needs a random sample of it.

Worked example: A Revision App and Exam Passes, First Compared by Who Chose to Use It and Then Tested by Random Allocation

Question A school looks at last year's results. Of the 120 students who chose to use a revision app, 84 passed their mathematics exam; of the 180 students who did not use it, 99 passed. (a) Is this an observational study or an experiment? Find the pass rate of each group and the difference between them. (b) Name a confounding variable that could explain part of the difference. This year, 200 volunteers are split at random into two groups of 100: the group given the app has 66 passes and the other group has 60. Find the difference now, and say what it suggests.

  1. 1.The students chose for themselves whether to use the app, and the school only recorded what happened. This is an observational study.

    last year: the students choseused app84 passed3670%no app99 passed8155%the school only recorded choicesan observational study
    last year: the students choseused app84 passed3670%no app99 passed8155%the school only recorded choicesan observational study
    The students decided for themselves whether to use the app, so this is an observational study.
  2. 2.(a) The pass rates are 84120 = 70% for the students who used the app and 99180 = 55% for the others, a difference of 15 percentage points. An observational study shows that using the app and passing go together; it cannot show that the app caused the passes.

    last year: the students choseused app84 passed3670%no app99 passed8155%70% − 55% = 15 points
    last year: the students choseused app84 passed3670%no app99 passed8155%70% − 55% = 15 points
    (a) The pass rates are 70% and 55%, a difference of 15 percentage points.
  3. 3.A confounding variable affects both the choice and the result. Students who are more motivated are more likely to download a revision app, and they are also more likely to pass whether or not they use it. Motivation, or the time a student spends revising, is a confounding variable here.

    the keener students chose the appused app84 passed3670%no app99 passed8155%motivation affects the choiceand the result: a confounder
    the keener students chose the appused app84 passed3670%no app99 passed8155%motivation affects the choiceand the result: a confounder
    Motivated students are more likely both to choose the app and to pass, so motivation is a confounding variable.
  4. 4.In the new study a random draw decides who gets the app, so motivated students are as likely to be in one group as in the other. The pass rates are 66100 = 66% with the app and 60100 = 60% without it.

    this year: 100 in each group, by random drawgiven app66 passed3466%not given60 passed4060%a random draw picks the groups66% and 60%
    this year: 100 in each group, by random drawgiven app66 passed3466%not given60 passed4060%a random draw picks the groups66% and 60%
    With random allocation, motivated students are spread evenly between the two groups.
  5. 5.(b) The difference is now 66% − 60% = 6 percentage points. Most of last year's 15 points therefore came from which students chose the app, not from the app itself. What is left is small: on two groups of 100, a gap of 6 points could arise by chance alone, so this study shows at most a modest effect.

    this year: 100 in each group, by random drawgiven app66 passed3466%not given60 passed4060%66% − 60% = 6 pointsmuch less than 15
    this year: 100 in each group, by random drawgiven app66 passed3466%not given60 passed4060%66% − 60% = 6 pointsmuch less than 15
    (b) The difference is now 6 percentage points: the app helps, but most of the first gap came from who chose it.

Answer: (a) an observational study; pass rates of 70% and 55%, a difference of 15 percentage points; (b) motivation, or the time spent revising, is a confounding variable; with random allocation the difference is 6 percentage points, so most of the first study's gap came from who chose the app, and a gap that size on 100 students each could be chance

Common mistakes

  • Reading the 15-point gap as the effect of the app. The students who chose it were not like the students who did not, so the gap mixes the effect of the app with the effect of motivation.
  • Letting the students, or their teachers, pick the groups in the new study. Only random allocation makes the two groups alike in everything except the app.

More sampling problems, worked step by step →

Practice Observational Studies and Experiments in the app