Distributions

Stage 16 of 23 Strand 2 of 5 16 lessons

16 illustrated lessons, each teaching the why before the how.

Revise Distributions with flashcards →

Jump to a lesson

Random Variables

A number attached to every outcome.

A random variable attaches a number to every outcome of an experiment

Roll a die and X is the score. Each value has its own probability.

Discrete means you can list the values: 0, 1, or 2 heads, with nothing in between.

Continuous means any value in a range, like a height, so its distribution is a curve, not bars.

Now you

Is the time a bus takes to arrive discrete or continuous?

Is the number of heads in 10 tosses discrete or continuous?

Probability Distributions

Every value with its chance, totaling one.

A probability distribution lists every value with its chance, and they must total one

A distribution lists each outcome and how likely that outcome is.

The probabilities must total 1, so a missing one is 1 minus the rest.

D is blank, so subtract the rest: 1 − 2/10 − 3/10 − 4/10 = 1/10.

Now you

Outcomes A, B and C have probabilities 4/10, 1/10 and 2/10. What is D?

Outcomes A, B and C have probabilities 4/10, 2/10 and 1/10. What is D?

Expected Value

The long-run average, weighted by chance.

The expected value is the long-run average, found by weighting each value by its chance

The prize is 0, 10 or 50, with probabilities 1/2, 2/5 and 1/10.

Weight each value by its probability and add: 0 + 4 + 5 = 9. That is E(X).

E(X) = 9 lands between the prizes, just short of 10. Nobody ever wins 9.

Charge 9 to play and the expected gain is 0. A game with E = 0 is called fair.

Now you

X is 0 with probability 1/2 and 14 with probability 1/2. What is E(X)?

X is 0 with probability 1/2 and 6 with probability 1/2. What is E(X)?

Variance

Average squared distance from the mean.

Variance measures spread by averaging the squared distance from the mean

Two variables can have the same mean and still be spread very differently.

This one also has mean 5, but its values spread right across the range.

Square each distance from the mean so they cannot cancel, then average.

X is 4 or 6 with probability 1/2 each: mean 5, squared distances 1, so Var(X) = 1.

Now you

X is 7 or 15, each with probability 1/2. What is Var(X)?

X is 5 or 11, each with probability 1/2. What is Var(X)?

Transforming a Random Variable

Adding moves the mean and leaves the spread.

Adding a constant shifts the mean but leaves the spread alone

Add 2 to every value and the whole distribution slides 2 to the right.

The mean slides with it: E(X + 2) = E(X) + 2.

Doubling doubles every distance from the mean, and squaring turns that 2 into 4.

A shift leaves the spread alone; doubling quadruples it: Var(2X) = 4Var(X).

Now you

E(X) = 3. What is E(3X + 5)?

Var(X) = 6. What is Var(2X + 9)?

Bernoulli Trials

Two outcomes, fixed chance, no memory.

A Bernoulli trial has two outcomes and a fixed probability, and each trial is independent of the last

One trial has two outcomes. Call them success and failure.

Repeat the trial and p never changes. Each trial is independent of the last.

Those three assumptions are what every binomial answer rests on.

Now you

Are these Bernoulli trials: rolling a die 10 times and counting sixes?

Are these Bernoulli trials: tossing the same coin 20 times?

The Binomial Distribution

Counting successes in n independent trials.

The binomial counts successes in a fixed number of independent trials

A binomial counts the successes in n independent trials, each with probability p.

X ~ B(n, p) names that count: n trials, each with probability p.

Choose which r trials succeed and multiply p for each success, 1 − p for each failure.

The nCr factor is why the bars follow a row of Pascal’s triangle.

Now you

X ~ B(4, 1/2). What is P(X = 2)?

In how many ways can exactly 2 of 5 trials succeed?

Binomial Mean and Variance

np, and np times one minus p.

A binomial has mean np and variance np(1 − p)

Take 5 trials with probability 1/2. The mean is np = 2.5, right under the peak.

The variance is np(1 − p): for these five trials, 5 × 1/2 × 1/2 = 1.25.

Take p down to 1/5 and the bars crowd to the left: np(1 − p) falls to 0.8.

At p = 1 every trial succeeds: one bar holds everything, and np(1 − p) is 0.

Now you

X ~ B(30, 1/5). What is the mean?

X ~ B(70, 1/10). What is the mean?

Modeling with a Binomial

Only as good as its quiet assumptions.

A binomial model is only as good as the assumptions behind it

Two percent of parts are faulty. In a box of 50, how many fail?

Model the count as a binomial: the expected number faulty is np = 1 per box.

If a bad batch makes faults come in clusters, the trials are not independent and the model no longer holds.

Now you

Can a binomial model be used if the sample is drawn without replacement from a small group?

Can a binomial model be used if p changes partway through?

Continuous Variables

Probability becomes area under a curve.

For a continuous variable probability is area, so a single exact value has none

A continuous variable has a curve instead of bars. The curve is called a density.

The whole area underneath is 1, exactly as the bars summed to 1.

A probability is the area over a range of values, so one exact value has none.

Now you

For a continuous variable, what is P(X = 4) exactly?

For a continuous variable, what is P(X = 2) exactly?

The Normal Distribution

The shape a binomial settles into.

The normal curve is symmetric about its mean and its spread is set by sigma

Increase the number of trials in a binomial and the bars settle into one shape.

That shape is the normal curve: symmetric, with one peak at the mean.

About 68 percent lies within one standard deviation, sigma, of the mean, either side.

Widen the band to two sigmas either side and it holds about 95 percent.

Three sigmas hold about 99.7 percent, so almost nothing is left outside.

Now you

Roughly what percent lies within 3 standard deviations of the mean?

Roughly what percent lies within 1 standard deviation of the mean?

Standardizing

How many sigmas from the mean a value sits.

Standardizing rewrites any value as the number of standard deviations from the mean

Subtract the mean from x to center it, then divide by sigma: z = (x − μ)/σ.

Now every normal curve becomes the same standard normal curve, with mean 0 and sigma 1.

A z of 1.5 means one and a half standard deviations above the mean.

Now you

The mean is 35 and sigma is 10. What is the z-score of 55?

The mean is 54 and sigma is 10. What is the z-score of 34?

Reading the Z-Table

Turn an area under the curve into a z-value.

A z-table turns an area under the normal curve into the z-value that cuts it off

A z-table gives the area to the left of a z-value — everything below it.

Each row pairs a z with the area beneath the curve to its left.

Want 5 percent in the upper tail? Then 95 percent lies to the left. Look up 0.9500.

That is where 1.645 and 1.96 come from: they are read from the table, not memorized.

The table lists positive z only; by symmetry −1.645 cuts the same 5 percent below.

Now you

The area to the left of z is 0.9750. What is z?

You want 1 percent in the upper tail. Which area do you look up?

Reading Normal Probabilities

The area to the left of a z-score.

Reading a normal probability means finding the area to the left of a z-score

Tables give the area to the left of a z-score. At z = 0 that is a half.

For the right-hand tail, subtract that area from 1.

For a band between two values, subtract the smaller area from the larger.

Now you

What is P(2.33 < Z < 2.58)?

What is P(1.645 < Z < 1.96)?

Inverse Normal

Given the area, find the cut-off value.

An inverse normal question gives you the area and asks for the value that cuts it off

The top 10 percent starts where 90 percent of the area sits to the left.

Read the table backward: which z has area 0.90 to its left? The table gives z = 1.28.

Then undo the standardizing to turn that z back into a value on the original scale.

Now you

The mean is 41 and sigma is 5. Which value has z = 2?

The mean is 25 and sigma is 2. Which value has z = -2?

Combining Normals

Means always add; variances only if independent.

Adding independent normal variables adds their means and adds their variances

Means add: E(X + Y) = E(X) + E(Y). With means 1 and 3, the sum has mean 4.

Var(X + Y) = Var(X) + Var(Y) = 1 + 1 = 2, so the sum is the wider curve.

Var(X − Y) = 1 + 1 = 2 as well — independent spreads add even when you subtract.

Means always add. Variances add only when X and Y are independent.

Now you

X and Y are independent with Var(X) = 6 and Var(Y) = 4. What is Var(X − Y)?

X and Y are independent with Var(X) = 8 and Var(Y) = 7. What is Var(X − Y)?

Continue your journey in the app — save your progress