A rare illness and a good test
1 person in 1000 has an illness. A test for it is right 99 times in 100: it says positive for 99 in 100 people who are ill, and negative for 99 in 100 people who are well. So it wrongly says positive for 1 in 100 well people.
Someone tests positive. Most people guess that they are ill with a chance of about 99 in 100. The true chance is about 1 in 11.
Out of 1000 people, 1 is inside the circle of people who are ill and 999 are outside it.
Count a thousand people
Think of 1000 people being tested. 1 of them is ill and 999 are well.
The ill person tests positive 99 times in 100, so count that person as a positive. The test wrongly says positive for 1 in 100 of the 999 well people: 999 ÷ 100 = 9.99, about 10 people. The other 989 well people test negative.
The 1000 people split into 1 who is ill and 999 who are well, and then by whether the test says yes, a positive, or no. The ill person tests positive. Of the well people, about 10 test positive, on the colored path, and 989 test negative.
Out of the positives
Everyone who tests positive is on one of two paths: ill and positive, or well and positive. That makes 1 + 10 = 11 people, and only 1 of them is ill. So the chance that a person with a positive test is ill is about , roughly 9 in 100. The other 10 positives are false alarms.
The test is not a bad test. Its mistakes are rare, but they happen among the 999 well people, while it can only catch the 1 ill person. A small fraction of a large group outnumbers a large fraction of a small one.
The 11 people who test positive, one part each. The shaded part is the 1 who is ill.
The exact counts
Counting 1000 people rounded 9.99 to 10 and 0.99 to 1. Take 100,000 people to keep every count whole. 1 in 1000 of them, 100 people, are ill, and 99,900 are well.
Of the 100 ill people, 99 test positive and 1 tests negative. Of the 99,900 well people, 1 in 100, which is 999, test positive, and 98,901 test negative. So 99 + 999 = 1,098 people test positive, and 99 of them are ill. The chance of being ill given a positive test is , about 0.090, or 9 in 100: the same as the rough count, .
100,000 people sorted by whether they are ill and whether the test says yes. The yes column holds 99 ill people and 999 well people, 1,098 in total.
The base rate decides
The fraction of people who have the illness before anyone is tested is called the base rate. Here it is 1 in 1000. What a positive result means depends on the base rate as much as on the test.
Take a different test that is right 9 times in 10, and use it on two groups of 100 people. In the first group, 10 people are ill. The test says positive for 9 of them, and for 1 in 10 of the 90 well people, which is 9 more. Of the 18 positives, 9 are ill, so a positive means ill with a chance of .
Each cell is one person. The 10 ill people are marked with a dot, and the 18 colored cells are the positives: 9 ill and 9 well.
In the second group, 50 people are ill. The test says positive for 45 of them, and for 1 in 10 of the 50 well people, which is 5. Of the 50 positives, 45 are ill, so a positive means ill with a chance of .
The test is the same in both groups. When the illness is commoner, more of the positives are true; when it is rarer, more of them are false.
The same test on 100 people with 50 ill, marked with a dot. Of the 50 colored positives, 45 are ill and 5 are well.
Testing again
This is why a positive result is usually followed by a second test. Among the 1,098 people who tested positive, 99 are ill: the base rate for the second test is about 1 in 11, not 1 in 1000.
Suppose the second test makes its mistakes independently of the first. It says positive again for 99 in 100 of the 99 ill people, about 98, and for 1 in 100 of the 999 well people, about 10. Of the 108 people with two positive tests, 98 are ill, so the chance of being ill is now about , roughly 9 in 10.
The usual mistakes
Answering 99 in 100. That is the chance that an ill person tests positive. The question asks the reverse: the chance that a person who tests positive is ill, out of everyone who tests positive.
Counting only the ill person, so about 1 person tests positive out of 1000. About 1 in 100 of the 999 well people test positive too, which is about 10 more.
Reading the 99 as a number of people. It is how often the test is right, 99 times in 100, not how many people test positive.
Saying that a rarer illness changes nothing because the test is the same. The test is the same, but fewer people are ill, so there are fewer true positives beside the same false ones.
A fraud flag
In the application below, a bank's flag checks card payments for fraud. Count 1000 payments, separate the true flags from the false ones, and divide the true flags by all the flags. Part (b) uses the same flag where fraud is ten times as common.
Worked example: A Bank's Fraud Flag, and the Chance That a Flagged Payment Is Really Fraud
Question A bank's fraud flag checks card payments. 1 in every 100 payments is fraudulent. The flag catches 9 out of every 10 fraudulent payments, but it also wrongly flags 1 in every 30 genuine payments. (a) Out of 1000 payments, how many are flagged? Find the probability that a flagged payment is really fraudulent. (b) At an online shop, 1 in every 10 payments is fraudulent, and the same flag is used. Find the probability that a flagged payment there is really fraudulent.
1.Of 1000 payments, 1100 of 1000, which is 10, are fraudulent, and 990 are genuine.
Each dot is one of 1000 payments. 1100 of them, 10 payments, are fraudulent. 2.The flag catches 910 of the 10 fraudulent payments, which is 9. It wrongly flags 130 of the 990 genuine payments, which is 33.
The flag catches 9 of the 10 fraudulent payments and wrongly flags 99030 = 33 genuine ones. 3.Altogether 9 + 33 = 42 payments are flagged. The false flags outnumber the true ones, because there are 99 genuine payments for every fraudulent one.
9 + 33 = 42 payments are flagged, and most of them are genuine, because fraud is rare. 4.(a) 42 payments are flagged, and P(fraudulent | flagged) = 942 = 314, so fewer than 1 flagged payment in 4 is really fraud.
(a) 42 payments are flagged, and P(fraudulent | flagged) = 942 = 314. 5.At the online shop, 1000 payments hold 100 fraudulent and 900 genuine ones. The flag catches 910 × 100 = 90 and wrongly flags 130 × 900 = 30, so 120 are flagged.
At the shop 100 of the 1000 payments are fraudulent: 90 are caught and 90030 = 30 genuine payments are flagged. 6.(b) P(fraudulent | flagged) = 90120 = 34. The flag works exactly as before; only the base rate of fraud has changed.
(b) 90120 = 34. The flag is the same; only the base rate has changed.
Answer: (a) 42 payments are flagged, and the probability is 314; (b) 34
Common mistakes
- Taking 910 as the probability that a flagged payment is fraudulent. 910 is the chance that a fraudulent payment is flagged, which is a different question with a different whole: the 10 fraudulent payments, not the 42 flagged ones.
- Leaving out the genuine payments that are flagged. They are a small fraction, 130, but of a large number, 990, so they give 33 flags, far more than the 9 from fraud.
More probability with several events problems, worked step by step →