Faulty parts in a box
A factory makes small parts, and 2% of them are faulty. The parts are packed 50 to a box. How many faulty parts will a box hold?
No single number answers that: one box may have none, another 1, another 3. What can be found is the probability of each count. Checking one part is a trial with two outcomes, faulty or fine, and a box is 50 trials. That is the shape of a binomial distribution, so the count X of faulty parts in a box might be modeled as X ~ B(50, 0.02).
Four conditions
A binomial model holds only when four conditions hold: a fixed number of trials, two outcomes on each trial, the same probability of success on every trial, and trials that are independent, so that the outcome of one does not change the chances of another.
For the box, the first two are certain: there are 50 parts, and each is faulty or fine. The last two are assumptions about the factory. Each part must be faulty with the same probability, 0.02, and one faulty part must say nothing about the next. The model answers the question only as well as those two assumptions are true.
What the model predicts
The expected count is np = 50 × 0.02 = 1 faulty part per box, and the variance is np(1 − p) = 50 × 0.02 × 0.98 = 0.98.
The probability of a box with no faulty part is , since all 50 parts must be fine. Exactly one is : the faulty part can be any of the 50. Exactly two is .
So the model says about 36% of boxes have no faulty part, and 1 − 0.3642 = 0.6358, about 64%, have at least one. Boxes with 5 or more faulty parts are rare: 0.0032, about 1 box in 300.
The probabilities of B(50, 0.02) for the number of faulty parts in a box. Almost all of the probability is on 0, 1 and 2, and the mean, 1, is at the peak.
When faults come in clusters
Suppose the parts come from two batches. In a good batch no part is faulty, and in a bad batch 4% are. Half the boxes are filled from each. Over all the boxes, 2% of parts are still faulty, and the mean is still 1 faulty part per box: 0 in a good box and 50 × 0.04 = 2 in a bad one, on average.
But the parts in one box are no longer independent. A faulty part means the box came from the bad batch, and then more faulty parts are likely. One fault makes the next more likely, so the faults cluster in the bad boxes.
Now a box has no faulty part with probability : every good box, and the bad boxes that happen to have none. The binomial model said 0.3642. A box with 5 or more faulty parts has probability 0.0245, nearly 8 times the model’s 0.0032, and the variance is 1.96, twice the model’s 0.98. The mean is right, and the model is wrong about everything else.
The clustered boxes on the same scale. The mean is still 1, but more than half of the boxes have no faulty part, and boxes with 3 or more are about twice as common as the binomial model says: 0.1616 against 0.0784.
Drawing without replacement
An inspector takes 5 parts from a box of 20 in which 2 are faulty. The first part is faulty with probability . If it is faulty, only 1 of the 19 left is; if it is fine, 2 of the 19 are. Each draw changes the next one’s chance, so the draws are not independent.
The binomial model B(5, 0.1) gives for no faulty part among the 5. The true value is .
Take the 5 parts from a batch of 2000 in which 200 are faulty, and the true value is 0.5902, almost exactly the binomial’s 0.5905. Removing 5 parts out of 2000 hardly changes the chance for the next one. So drawing without replacement breaks the model when the sample is a large part of the group, and barely matters when it is a small part.
Checking a situation
Ask each of the four questions in turn. Is the number of trials fixed? Does each trial have two outcomes? Is the probability the same on every trial? Are the trials independent?
Independent trials with a fixed n and a fixed p are exactly what the model needs, so it applies. If a machine wears during the day and its fault rate rises from 2% to 5%, p changes partway through and the model fails. Making n larger does not repair a broken condition: the fixed p and the independence are needed at every n.
Worked example: A Guessed Multiple-Choice Quiz That Is Binomial, and Question Cards Drawn from a Box That Are Not
Question A quiz has 5 questions, each with 4 options, and a student guesses every answer. (a) Explain why X, the number she gets right, can be modeled by B(5, 14), and find the probability that she gets at least 4 right. (b) In a later round the quiz-master draws 5 question cards at random, without replacement, from a box of 12 in which 3 are on sport. A friend models S, the number of sport questions, by B(5, 14). Explain why this model fails, and find P(S = 0) correctly and by the friend's model.
1.There is a fixed number of trials, 5 questions. Each has two outcomes, right or wrong. The guesses are independent of one another, and each is right with the same probability 14. So X ∼ B(5, 14).
The guesses meet all four conditions, so X ∼ B(5, 14). 2.Exactly 4 right, with the one wrong answer on any of the 5 questions: P(X = 4) = 54 (14)4 × 34 = 151024. All 5 right: P(X = 5) = (14)5 = 11024.
The bars are written in 1024ths: 45 = 1024 equally likely answer sheets. Four right happens on 15 of them and five right on 1. 3.(a) P(X ≥ 4) = 151024 + 11024 = 161024 = 164, about 0.0156.
(a) P(X ≥ 4) = 161024 = 164. 4.The cards are drawn without replacement. After a sport card, only 2 of the 11 cards left are on sport; after a card on another subject, 3 of the 11 are. The chance on each draw depends on the draws before it, so the draws are not independent and the binomial model fails.
The cards are not put back. After a sport card 2 of the 11 left are on sport; after any other card 3 are. The draws are not independent. 5.Correctly, all 5 cards come from the 9 that are not on sport: P(S = 0) = 95125 = 126792 = 744 ≈ 0.159.
Correctly, P(S = 0) = 95125 = 126792 = 744. 6.(b) The friend's model gives P(S = 0) = (34)5 = 2431024 ≈ 0.237, about half as large again as the true 0.159. The model fails because the draws are not independent.
(b) The binomial model gives 2431024 ≈ 0.237 against the true 0.159: it fails because the draws are not independent.
Answer: (a) 164; (b) the draws are not independent, since each card drawn changes what is left in the box; P(S = 0) = 744 ≈ 0.159, against 2431024 ≈ 0.237 by the model
Common mistakes
- Saying the model fails because p is not 14. Taken on its own, any one card is on sport with probability 312 = 14; what fails is independence, since each draw changes what the next one can be.
- Leaving out 54 in P(X = 4) and writing 31024. The one wrong answer can be on any of the 5 questions, so there are 5 ways to get exactly 4 right.
More probability distributions problems, worked step by step →