One trial, two outcomes
Roll a dice once and ask one question: is it a six? There are only two answers. A trial with exactly two outcomes is called a Bernoulli trial. The outcome being counted is called a success and the other a failure, whether or not it is good news: a faulty part on a production line can be the success.
The probability of a success is written p, so the probability of a failure is 1 − p. For a six on a fair dice, and . For a player who scores 70% of free throws, one throw is a Bernoulli trial with p = 0.7.
One Bernoulli trial: a success with probability p, or a failure with probability 1 − p. The two branches add up to 1.
One trial as a random variable
Let X be 1 for a success and 0 for a failure. Its mean is E(X) = 1 × p + 0 × (1 − p) = p.
Because and , as well. So Var. For a six on a dice, the mean is and the variance is .
Repeating the trial
Roll the dice again. A run of Bernoulli trials needs two more things. The probability of success must be the same p on every trial: the second roll gives a six with probability , just as the first did. And the trials must be independent: the result of the first roll does not change the chances on the second.
With both in place, the probability of a run of results is a product. Two rolls give four outcomes: a six then a six with probability , a six then not a six with , not a six then a six with , and two failures with . Check: 1 + 5 + 5 + 25 = 36, so the four add up to .
Two trials, S for success and F for failure. The second pair of branches is p and 1 − p after a success and after a failure alike: that is independence with a fixed p. Each outcome is the product along its path.
The conditions for a binomial count
Counting the successes in a run of trials leads to the binomial distribution. It rests on four conditions: a fixed number of trials n, two outcomes on each trial, the same probability p on every trial, and trials that are independent of each other.
All four must hold. Tossing the same coin 20 times meets them: 20 trials, heads or tails, each time, and no toss affects another. Rolling a dice 10 times and counting sixes meets them too, with .
The number of trials does not decide it. A small n can meet every condition, and a large n can still fail one.
When a condition fails
Draw two cards from a pack of 52 without putting the first one back, and count the hearts. The first card is a heart with probability . But the second card depends on the first. After a heart, 12 of the 51 cards left are hearts, so the chance is . After another suit, 13 of the 51 are hearts, so it is .
The probability changes with what happened before, so the draws are not independent, and they do not form a run of Bernoulli trials. If the first card is put back and the pack shuffled, every draw is a heart with probability again, and the conditions hold.
Two cards drawn without replacement. The second pair of branches is different after a heart and after another suit , so the second draw depends on the first.
An application
In the application below, the conditions are checked twice: once for a student who guesses every answer on a quiz, where they hold, and once for question cards drawn from a box without replacement, where they fail.
Worked example: A Guessed Multiple-Choice Quiz That Is Binomial, and Question Cards Drawn from a Box That Are Not
Question A quiz has 5 questions, each with 4 options, and a student guesses every answer. (a) Explain why X, the number she gets right, can be modeled by B(5, 14), and find the probability that she gets at least 4 right. (b) In a later round the quiz-master draws 5 question cards at random, without replacement, from a box of 12 in which 3 are on sport. A friend models S, the number of sport questions, by B(5, 14). Explain why this model fails, and find P(S = 0) correctly and by the friend's model.
1.There is a fixed number of trials, 5 questions. Each has two outcomes, right or wrong. The guesses are independent of one another, and each is right with the same probability 14. So X ∼ B(5, 14).
The guesses meet all four conditions, so X ∼ B(5, 14). 2.Exactly 4 right, with the one wrong answer on any of the 5 questions: P(X = 4) = 54 (14)4 × 34 = 151024. All 5 right: P(X = 5) = (14)5 = 11024.
The bars are written in 1024ths: 45 = 1024 equally likely answer sheets. Four right happens on 15 of them and five right on 1. 3.(a) P(X ≥ 4) = 151024 + 11024 = 161024 = 164, about 0.0156.
(a) P(X ≥ 4) = 161024 = 164. 4.The cards are drawn without replacement. After a sport card, only 2 of the 11 cards left are on sport; after a card on another subject, 3 of the 11 are. The chance on each draw depends on the draws before it, so the draws are not independent and the binomial model fails.
The cards are not put back. After a sport card 2 of the 11 left are on sport; after any other card 3 are. The draws are not independent. 5.Correctly, all 5 cards come from the 9 that are not on sport: P(S = 0) = 95125 = 126792 = 744 ≈ 0.159.
Correctly, P(S = 0) = 95125 = 126792 = 744. 6.(b) The friend's model gives P(S = 0) = (34)5 = 2431024 ≈ 0.237, about half as large again as the true 0.159. The model fails because the draws are not independent.
(b) The binomial model gives 2431024 ≈ 0.237 against the true 0.159: it fails because the draws are not independent.
Answer: (a) 164; (b) the draws are not independent, since each card drawn changes what is left in the box; P(S = 0) = 744 ≈ 0.159, against 2431024 ≈ 0.237 by the model
Common mistakes
- Saying the model fails because p is not 14. Taken on its own, any one card is on sport with probability 312 = 14; what fails is independence, since each draw changes what the next one can be.
- Leaving out 54 in P(X = 4) and writing 31024. The one wrong answer can be on any of the 5 questions, so there are 5 ways to get exactly 4 right.
More probability distributions problems, worked step by step →