Statistical Inference · applications

Applications: Statistical Inference

10 question types · Pre-University · each worked step by step with a figure that follows the steps

H2

01

Strawberries Weighed from a Day's Picking, and the Mean and Spread of the Whole Crop Estimated from Two Sums

methodTake the Sample Mean for the Mean; for the Variance, Divide the Sum of Squares About the Mean by n − 1, Not by n

A fruit farm picks 25 strawberries at random from a day's crop and weighs them. The masses x grams give ∑ x = 605 and ∑ x2 = 14941. (a) Find unbiased estimates of the mean and the variance of the mass of a strawberry in the crop. (b) A worker divides the sum of squares about the mean by 25 instead of 24. Find the value this gives, say why it tends to be too small, and use your unbiased estimate from (a) to find the standard error of the sample mean.

25 berries605 g in allmean = 605/25 = 24.2 g
The 25 masses add up to 605 g, so the sample mean is x = 24.2 g.
The unbiased estimate of the population mean is the sample mean: x = 60525 = 24.2 g.
step 1 of 5

A point estimate is a single number, worked out from a sample, that stands for a quantity of the whole population. An estimate is unbiased when its average over all possible samples equals the true value. The sample mean is an unbiased estimate of the population mean. For the variance, the squared deviations are measured from the sample's own mean, which makes them a little too small, so their sum is divided by n − 1 rather than n.

  1. The unbiased estimate of the population mean is the sample mean: x = 60525 = 24.2 g.
  2. The sum of squares about the mean is Sxx = ∑ x2 − (∑ x)2n = 14941 − 605225 = 14941 − 14641 = 300.
  3. (a) The unbiased estimate of the variance divides by n − 1 = 24: s2 = 30024 = 12.5. The estimates are 24.2 g for the mean and 12.5 g2 for the variance.
  4. Dividing by 25 gives 30025 = 12. The deviations are measured from x, the center of this sample itself, so their squares are smaller on average than squares measured from the true mean μ. Dividing by n − 1 instead of n makes up for that.
  5. (b) The worker's value is 12 g2, which tends to underestimate the variance. The standard error of the mean is estimated as √s2n = √12.525 = √0.5 = 0.707 g. Check: 24 × 12.5 = 25 × 12 = 300.

answer(a) the mean 24.2 g and the variance 12.5 g2; (b) 12 g2, too small because the deviations are measured from the sample's own mean; the standard error is about 0.707 g

techniqueUnbiased Sample Variance · Point Estimates

examsH2

Common pitfalls

  • Working out ∑ x2n − x2 = 597.64 − 585.64 = 12 and calling it the unbiased estimate. That formula divides by n; the unbiased estimate is nn − 1 times it, 2524 × 12 = 12.5.
  • Subtracting 6052 itself from 14941, or subtracting 60525 without squaring it. The square of the total must be divided by n: 605225 = 14641 is 25 times the square of the mean.
02

A Coffee Machine Checked After a Repair, and a Two-Tailed Test of Whether Its Mean Pour Has Changed

methodStandardize the Sample Mean by the Standard Deviation over the Square Root of n, Double the Tail for a Two-Tailed p-Value, and Compare It with the Level

A coffee machine is set to pour 250 ml into each cup, and the volume it pours is normally distributed with standard deviation 6 ml. After a repair, a sample of 36 cups has a mean volume of 252.1 ml. (a) Test at the 5% level whether the mean volume has changed: state the hypotheses, and find the test statistic and the p-value. (b) Would the conclusion be the same at the 1% level? Give both conclusions in context.

0Z ~ N(0, 1) if H0 is trueH0: mean 250 mlH1: mean not 250 ml, two tails
If H0 is true, the standardized sample mean follows N(0, 1). A change in either direction counts, so both tails matter.
Let μ be the mean volume after the repair. The question is whether it has changed in either direction, so the test is two-tailed: H0: μ = 250 and H1: μ ≠ 250.
step 1 of 5

A hypothesis test starts from a null hypothesis H0, here that nothing has changed, and asks how surprising the sample would be if H0 were true. The p-value is the probability, under H0, of a result at least as extreme as the one observed. When the question is whether a mean has changed, in either direction, the test is two-tailed and extreme results in both tails count.

  1. Let μ be the mean volume after the repair. The question is whether it has changed in either direction, so the test is two-tailed: H0: μ = 250 and H1: μ ≠ 250.
  2. Under H0 the mean of 36 cups is normal with mean 250 and standard deviation 6√36 = 1 ml. The test statistic is z = 252.1 − 2501 = 2.1.
  3. The p-value counts results at least this far from 250 in either direction: 2 × P(Z > 2.1) = 2 × (1 − 0.98214) = 0.0357.
  4. (a) 0.0357 < 0.05, or equally 2.1 > 1.96, so reject H0. There is evidence at the 5% level that the mean volume has changed from 250 ml since the repair.
  5. (b) At the 1% level the critical values are ± 2.576, and 0.0357 > 0.01, so do not reject H0. At 1% there is not enough evidence that the mean volume has changed. The same data can be significant at one level and not at another, which is why the level is chosen before the data are seen.

answer(a) H0: μ = 250, H1: μ ≠ 250; z = 2.1 and the p-value is 0.0357, below 0.05 (2.1 > 1.96): reject H0, there is evidence that the mean volume has changed; (b) no: 0.0357 > 0.01 (2.1 < 2.576), so at the 1% level there is not enough evidence of a change

techniqueThe Z-Test · P-Values · One-Tailed and Two-Tailed

examsH2

Common pitfalls

  • Standardizing with the standard deviation of one cup: 252.1 − 2506 = 0.35. The test is about the mean of 36 cups, whose standard deviation is 6√36 = 1 ml.
  • Reporting P(Z > 2.1) = 0.0179 as the p-value. The alternative is μ ≠ 250, so a result as far below 250 would count just as much, and the tail is doubled.
03

A Basketball Player's Free Throws After a Summer of Coaching, Tested Against Last Season's Rate

methodAdd Binomial Terms from the Top Until the Tail Would Pass the Level; the Last Tail Below It Is the Critical Region, and Its Probability Is the Actual Significance Level

Last season a basketball player scored 40% of her free throws. After a summer of coaching she believes she has improved, and she takes 20 free throws to find out. (a) Using a binomial model, find the critical region for a test at the 5% level of H0: p = 0.4 against H1: p > 0.4, and the actual significance level of the test. (b) She scores 12 of the 20. State the conclusion of the test in context.

05101520X ~ B(20, 0.4) if H0 is trueX ~ B(20, 0.4) under H0H1: p > 0.4, so the top tail
Under H0 the number she scores is X ∼ B(20, 0.4). The alternative p > 0.4 puts the critical region in the top tail.
Let X be the number of free throws she scores out of 20. Under H0, X ∼ B(20, 0.4). The alternative is p > 0.4, so the test is one-tailed and the critical region is at the top: X ≥ k for the smallest k with P(X ≥ k) ≤ 0.05.
step 1 of 5

The critical region is the set of results that would lead to rejecting H0. For a discrete distribution the probability of the region cannot usually be exactly 5%, so the region is the largest tail whose probability does not exceed 5%, and that probability is the actual significance level. A result outside the region means that H0 is not rejected, which is not the same as showing that it is true.

  1. Let X be the number of free throws she scores out of 20. Under H0, X ∼ B(20, 0.4). The alternative is p > 0.4, so the test is one-tailed and the critical region is at the top: X ≥ k for the smallest k with P(X ≥ k) ≤ 0.05.
  2. Add the binomial terms from the top: P(X ≥ 13) = 0.0210 and P(X ≥ 12) = 0.0565. The first is below 0.05 and the second is above it.
  3. (a) The critical region is X ≥ 13. The actual significance level is P(X ≥ 13) = 0.0210, about 2.1%: that is the chance of rejecting H0 when in fact she has not improved.
  4. She scored 12, which is not in the critical region. In the same way, P(X ≥ 12) = 0.0565 is more than 0.05.
  5. (b) Do not reject H0. There is not enough evidence at the 5% level that her success rate has risen above 40%. This does not show that she has not improved: 12 out of 20 is 60%, but 20 throws are too few to tell that rate apart from chance.

answer(a) the critical region is X ≥ 13, and the actual significance level is 0.0210; (b) 12 is not in the critical region, since P(X ≥ 12) = 0.0565, so do not reject H0: there is not enough evidence that her success rate has risen above 40%

techniqueCritical Regions · Testing a Proportion with the Binomial · Null and Alternative Hypotheses

examsH2

Common pitfalls

  • Using P(X = 13) = 0.0146, the chance of exactly 13, in place of P(X ≥ 13). The test asks how likely a result this extreme or more extreme is, so every term from 13 to 20 is added.
  • Taking X ≥ 12 as the critical region because 0.0565 is close to 5%. The region's probability must not exceed the level, so 12 is left out and the region starts at 13.
04

A Bakery's 400 g Loaves Inspected by Weighing 25, and the Chance That Light Loaves Pass the Test

methodFix the Critical Region Under H0, Then Find the Chance of Falling Outside It Under the True Mean: That Chance Is the Type II Error

A bakery's loaves are labeled 400 g, and their masses are normally distributed with standard deviation 10 g. An inspector weighs 25 loaves and will test H0: μ = 400 against H1: μ < 400 at the 5% level. (a) Find the critical region for the sample mean, and state the probability of a Type I error. (b) In fact the mean mass is 395 g. Find the probability that the test fails to detect this, and say what kind of error that is.

400if H0 is true: mean 400sd of the mean = 10/√25 = 2 gH1: mean < 400, the lower tail
If H0 is true, the mean of 25 loaves is normal with mean 400 g and standard deviation 2 g.
Under H0 the mean of 25 loaves is normal with mean 400 g and standard deviation 10√25 = 2 g. The alternative is μ < 400, so the test is one-tailed with the critical region in the lower tail.
step 1 of 5

A Type I error is rejecting H0 when it is true; its probability is the significance level. A Type II error is failing to reject H0 when it is false. Its probability depends on the true value, so it is found in two stages: first the critical region is fixed using H0, then the probability of a sample mean outside that region is worked out using the true mean.

  1. Under H0 the mean of 25 loaves is normal with mean 400 g and standard deviation 10√25 = 2 g. The alternative is μ < 400, so the test is one-tailed with the critical region in the lower tail.
  2. The lowest 5% of the standard normal distribution lies below z = −1.645, so the critical value is 400 − 1.645 × 2 = 396.71 g.
  3. (a) The critical region is x < 396.71 g. A Type I error is rejecting H0 when it is true, and its probability is the significance level, 0.05.
  4. With the true mean 395 g, the sample mean is normal with mean 395 and standard deviation 2. The test fails to reject H0 when x ≥ 396.71, which is z = 396.71 − 3952 = 0.855 standard deviations above 395.
  5. (b) P(x ≥ 396.71) = 1 − 0.8037 = 0.196. This is a Type II error, failing to reject H0 when it is false: about one inspection in five would pass loaves that are on average 5 g light. Weighing more loaves would make this probability smaller.

answer(a) the critical region is x < 396.71 g, and the probability of a Type I error is 0.05; (b) 0.196, the probability of a Type II error

techniqueType One and Type Two Errors · Critical Regions · One-Tailed and Two-Tailed

examsH2

Common pitfalls

  • Using the standard deviation of one loaf, 10 g, to find the critical value, which gives 400 − 16.45 = 383.55 g. The test uses the mean of 25 loaves, whose standard deviation is 2 g.
  • Answering that the probability of a Type II error is 1 − 0.05 = 0.95. That is the probability of not rejecting H0 when H0 is true; a Type II error happens when H0 is false, so it is worked out with the true mean of 395 g.
05

Journey Times on a Bus Route, a 95% Confidence Interval for the Mean, and the Timetable's 32 Minutes

methodGo 1.96 Standard Errors Either Side of the Sample Mean, Then Set the Timetable Against the Interval and Ask What the Interval Estimates

A bus company times 64 journeys on one route, chosen at random over a month. The mean time is 34.5 minutes, and the unbiased estimate of the standard deviation is 8 minutes. (a) Find a 95% confidence interval for the mean journey time on the route. (b) The timetable allows 32 minutes. Use the interval to comment on the timetable, and say what the interval does not tell you.

minutes34.5standard error = 8/√64 = 1 minute
The sample mean is 34.5 minutes, and its standard error is 8√64 = 1 minute.
The sample is large, so by the central limit theorem the sample mean is approximately normal, and s = 8 can be used in place of σ. The standard error is 8√64 = 1 minute.
step 1 of 5

A confidence interval is a range of plausible values for a population mean, built around the sample mean. A 95% interval reaches 1.96 standard errors either side, because 95% of a normal distribution lies within 1.96 standard deviations of its mean. The 95% describes the method: of many intervals made this way from different samples, about 95% contain the true mean.

  1. The sample is large, so by the central limit theorem the sample mean is approximately normal, and s = 8 can be used in place of σ. The standard error is 8√64 = 1 minute.
  2. A 95% interval reaches 1.96 standard errors either side of the sample mean: 34.5 ± 1.96 × 1 = 34.5 ± 1.96.
  3. (a) The 95% confidence interval is (32.54, 36.46) minutes.
  4. 32 minutes lies below the interval, so a mean of 32 minutes is not a plausible value at this level: the timetable allows too little time. A two-tailed test of μ = 32 at the 5% level would reject it.
  5. (b) The interval is about the mean journey time. It does not say that 95% of journeys take between 32.54 and 36.46 minutes: with a standard deviation of 8 minutes, single journeys vary far more than that. Nor is there a 95% chance that this one interval contains μ; the 95% describes how often intervals made this way contain the true mean.

answer(a) (32.54, 36.46) minutes; (b) 32 is below the interval, so the timetable allows too little time; the interval is for the mean, not for single journeys, and the 95% describes how often intervals made this way contain the true mean

techniqueConfidence Intervals · Point Estimates

examsSAT

Common pitfalls

  • Going 1.96 × 8 either side of the mean, which gives (18.82, 50.18). That range is for single journeys; the interval for the mean uses the standard error 8√64 = 1.
  • Saying that 95% of journeys take between 32.54 and 36.46 minutes. The interval estimates the mean of all journeys; a single journey is often more than 8 minutes from the mean.
06

Film Preferences of Cinema Customers in Two Age Groups, Tested for Independence

methodEach Expected Frequency Is the Row Total Times the Column Total Divided by the Grand Total; the Degrees of Freedom Are (Rows − 1) Times (Columns − 1)

A cinema asks 200 customers which kind of film they prefer. Of the 120 customers under 30, 64 choose action, 28 comedy and 28 drama; of the 80 aged 30 or over, 26 choose action, 32 comedy and 22 drama. (a) Find the expected frequencies if the preferred kind of film is independent of age group, and the value of the test statistic χ2. (b) Test at the 5% level whether preference and age group are independent.

actioncomedydramatotalunder 3064282812030 or over26322280total906050200H0: film and age group are independentE = row total × column total/200
The observed table, with its row and column totals in gold.
H0: the preferred kind of film is independent of age group. H1: it is not independent. The column totals are 64 + 26 = 90 for action, 28 + 32 = 60 for comedy and 28 + 22 = 50 for drama.
step 1 of 5

If two classifications are independent, the share of each row that falls in a column is the same for every row, so the expected frequency of a cell is its row total times its column total divided by the grand total. The statistic χ2 = ∑ (O − E)2E measures how far the observed table is from those expected frequencies. For a table of r rows and c columns it is compared with the χ2 distribution on (r − 1)(c − 1) degrees of freedom.

  1. H0: the preferred kind of film is independent of age group. H1: it is not independent. The column totals are 64 + 26 = 90 for action, 28 + 32 = 60 for comedy and 28 + 22 = 50 for drama.
  2. For customers under 30 the expected frequencies are 120 × 90200 = 54, 120 × 60200 = 36 and 120 × 50200 = 30; for customers aged 30 or over they are 36, 24 and 20. Every one is at least 5, so the test can be used.
  3. (a) χ2 = ∑ (O − E)2E = 10254 + 8236 + 2230 + 10236 + 8224 + 2220 = 9.407.
  4. The table has 2 rows and 3 columns, so there are (2 − 1)(3 − 1) = 2 degrees of freedom. The critical value of χ2(2) at the 5% level is 5.991.
  5. (b) 9.407 > 5.991, so reject H0. There is evidence at the 5% level that the kind of film customers prefer depends on their age group: customers under 30 choose action more often than independence would give, and older customers choose comedy more often. Check: in each row the differences O − E add up to 0, since 10 − 8 − 2 = 0.

answer(a) the expected frequencies are 54, 36 and 30 under 30, and 36, 24 and 20 at 30 or over; χ2 = 9.407; (b) 9.407 > 5.991 on 2 degrees of freedom, so reject H0: the kind of film preferred depends on age group

techniqueThe Chi-Squared Test for Independence · Expected Frequencies

Common pitfalls

  • Using 6 − 1 = 5 degrees of freedom, one fewer than the number of cells. Once the row and column totals are fixed, only 2 cells of a 2 × 3 table can be chosen freely, so there are (2 − 1)(3 − 1) = 2.
  • Dividing each (O − E)2 by the observed frequency O instead of the expected frequency E. The statistic measures the distance from what H0 predicts, so each term is divided by E.
07

A Games Club's Suspect Die Thrown 120 Times and Tested for Fairness

methodExpect One Sixth of the Throws for Each Score, Add (O − E) Squared over E, and Use One Degree of Freedom Fewer Than the Number of Classes

A games club suspects that one of its dice is not fair. It is thrown 120 times, and the scores 1 to 6 come up 14, 22, 18, 26, 17 and 23 times. (a) Find the expected frequencies if the die is fair, and the value of χ2. (b) Test at the 5% level whether the die is fair, and state the conclusion in context.

141222183264175236fair die: 20 of each scoreE = 120 × 1/6 = 20 for each score
The bars are the observed frequencies; the dashed line is the expected frequency of 20 for a fair die.
H0: the die is fair, so each score has probability 16. H1: the die is not fair. The expected frequency of each score is 120 × 16 = 20.
step 1 of 5

A goodness-of-fit test compares observed frequencies with the frequencies a stated distribution predicts. The statistic χ2 = ∑ (O − E)2E is compared with the χ2 distribution whose degrees of freedom are the number of classes minus the number of constraints; here the only constraint is that the expected frequencies add up to the same total as the observed ones.

  1. H0: the die is fair, so each score has probability 16. H1: the die is not fair. The expected frequency of each score is 120 × 16 = 20.
  2. The differences O − E are −6, 2, −2, 6, −3 and 3, which add up to 0 as they must. Their squares are 36, 4, 4, 36, 9 and 9, a total of 98.
  3. (a) Every expected frequency is 20, so χ2 = 9820 = 4.9.
  4. There are 6 classes and one constraint, that the frequencies add up to 120, so there are 6 − 1 = 5 degrees of freedom. The critical value of χ2(5) at the 5% level is 11.070.
  5. (b) 4.9 < 11.070, so do not reject H0. There is not enough evidence at the 5% level that the die is unfair: differences as large as these often happen with a fair die in 120 throws. Had the statistic been above 11.070, the club would have had evidence that the die is biased.

answer(a) each expected frequency is 20, and χ2 = 4.9; (b) 4.9 < 11.070 on 5 degrees of freedom, so do not reject H0: there is not enough evidence that the die is unfair

techniqueThe Chi-Squared Goodness of Fit Test · Expected Frequencies · Null and Alternative Hypotheses

examsH2

Common pitfalls

  • Using 6 degrees of freedom, one for each score. The six expected frequencies must add up to 120, which fixes the last one, so there are 5.
  • Concluding that the test has proved the die is fair. Not rejecting H0 only means that 120 throws give no clear evidence against fairness; a small bias could still be there.
08

Nine Batteries Tested Against a Manufacturer's Claim of 20 Hours

methodWith the Population Standard Deviation Unknown and a Small Sample, Estimate It by s and Compare the Statistic with the t Distribution on n − 1 Degrees of Freedom

A manufacturer claims that its batteries last 20 hours on average in a torch. A consumer group suspects they last less, and tests 9 batteries, which last 17.6, 18.0, 18.4, 18.6, 19.0, 19.4, 20.2, 20.4 and 21.2 hours. Battery lives are assumed to be normally distributed. (a) Find the sample mean, the unbiased estimate of the variance and the test statistic. (b) Test the claim at the 5% level, and say whether the conclusion would be the same at the 1% level.

17182122claim 20H0: mean 20, H1: mean < 20sd unknown, n = 9: use 8 df
The nine battery lives, in hours, against the claimed mean of 20.
Let μ be the mean life of the batteries. H0: μ = 20 and H1: μ < 20, a one-tailed test. The variance is unknown and the sample is small, so the test uses the t distribution with 9 − 1 = 8 degrees of freedom.
step 1 of 5

When the population standard deviation is unknown it is estimated by s from the sample. For a small sample from a normal population, x − μs / √n then follows the t distribution with n − 1 degrees of freedom, which has heavier tails than the normal distribution, so its critical values are further out.

  1. Let μ be the mean life of the batteries. H0: μ = 20 and H1: μ < 20, a one-tailed test. The variance is unknown and the sample is small, so the test uses the t distribution with 9 − 1 = 8 degrees of freedom.
  2. The nine lives add up to 172.8, so x = 172.89 = 19.2 hours. Their deviations from 19.2 are −1.6, −1.2, −0.8, −0.6, −0.2, 0.2, 1.0, 1.2 and 2.0 hours, and their squares add up to 11.52.
  3. (a) s2 = 11.528 = 1.44, so s = 1.2 hours, and the test statistic is 19.2 − 201.2 / √9 = −0.80.4 = −2.0.
  4. (b) For a lower-tailed test at 5% on 8 degrees of freedom the critical value is −1.860. Since −2.0 < −1.860, reject H0: there is evidence at the 5% level that the batteries last less than 20 hours on average.
  5. At the 1% level the critical value is −2.896, and −2.0 > −2.896, so H0 is not rejected: at 1% there is not enough evidence against the manufacturer's claim. Check: the deviations add up to 0, as deviations from a mean must.

answer(a) x = 19.2 hours, s2 = 1.44, and the test statistic is −2.0; (b) −2.0 < −1.860, so at 5% reject H0: there is evidence that the batteries last less than 20 hours on average; at 1% the critical value is −2.896 and H0 is not rejected

techniqueThe One-Sample T-Test · One-Tailed and Two-Tailed · Null and Alternative Hypotheses

examsH2

Common pitfalls

  • Comparing −2.0 with the normal critical value −1.645. With σ estimated from only 9 batteries the statistic has the heavier-tailed t(8) distribution, and its critical value is −1.860.
  • Dividing the sum of squares by 9, which gives 1.28, or dividing s by 9 instead of by √9 = 3. The unbiased variance divides by n − 1 = 8, and the standard error is s√n = 0.4.
09

Tomato Seedlings Grown with Two Fertilizers and Compared by Their Mean Heights

methodPool the Two Variance Estimates, Each Weighted by Its Degrees of Freedom, Then Divide the Difference in Means by Its Standard Error

A gardener grows 12 tomato seedlings, giving fertilizer A to 6 of them and fertilizer B to the other 6, chosen at random. After four weeks the heights in centimeters are: A: 22, 22, 24, 25, 25, 26; B: 19, 19, 21, 21, 23, 23. Heights are assumed to be normal, with the same variance for both fertilizers. (a) Find the pooled estimate of the variance and the test statistic for comparing the two means. (b) Test at the 5% level whether the fertilizers give different mean heights, and then at the 1% level.

1820222426ABmean 24mean 21A: mean 144/6 = 24 cmB: mean 126/6 = 21 cm
The heights of the seedlings, in centimeters, with each sample's mean.
Let μA and μB be the mean heights. H0: μA = μB and H1: μA ≠ μB, a two-tailed test. The sample means are xA = 1446 = 24 cm and xB = 1266 = 21 cm.
step 1 of 5

When two normal populations share one variance, the two sample estimates are combined into a pooled estimate, each weighted by its degrees of freedom. The difference in sample means, divided by its standard error √sp2 (1n1 + 1n2), follows the t distribution with n1 + n2 − 2 degrees of freedom when H0 is true.

  1. Let μA and μB be the mean heights. H0: μA = μB and H1: μA ≠ μB, a two-tailed test. The sample means are xA = 1446 = 24 cm and xB = 1266 = 21 cm.
  2. The squared deviations add up to 4 + 4 + 0 + 1 + 1 + 4 = 14 for A and 4 + 4 + 0 + 0 + 4 + 4 = 16 for B, so sA2 = 145 = 2.8 and sB2 = 165 = 3.2.
  3. The pooled variance weights each estimate by its degrees of freedom: sp2 = 5 × 2.8 + 5 × 3.210 = 3010 = 3.
  4. (a) The standard error of the difference is √3 × (16 + 16) = √1 = 1 cm, so the test statistic is 24 − 211 = 3.0, on 6 + 6 − 2 = 10 degrees of freedom.
  5. (b) At 5%, two-tailed, the critical value of t(10) is 2.228, and 3.0 > 2.228: reject H0. There is evidence that the fertilizers give different mean heights, with the seedlings given A taller. At 1% the critical value is 3.169, and 3.0 < 3.169, so at that level the difference is not significant.

answer(a) the pooled variance is 3 and the test statistic is 3.0, on 10 degrees of freedom; (b) 3.0 > 2.228, so at 5% reject H0: the fertilizers give different mean heights; at 1% the critical value is 3.169, so there the difference is not significant

techniqueThe Two-Sample T-Test

Common pitfalls

  • Using 6 − 1 = 5 degrees of freedom, as for one sample. The pooled estimate uses both samples, which gives 5 + 5 = 10 degrees of freedom.
  • Adding the two variances without weighting or dividing, 2.8 + 3.2 = 6, and using √66. The pooled variance is the weighted average 3, and it is multiplied by 16 + 16.
10

Bolts from Two Machines, One Tested Against the Specification's Variance and Then Against the Other Machine

methodFor One Variance, Compare (n − 1) Times s Squared over the Stated Variance with Chi-Squared; for Two, Put the Larger Estimate over the Smaller and Use F

A factory's specification says that the lengths of its bolts must have a variance of no more than 0.04 mm2. A sample of 11 bolts from machine A gives an unbiased variance estimate sA2 = 0.08 mm2, and a sample of 13 bolts from machine B gives sB2 = 0.032 mm2. Lengths are normally distributed. (a) Test at the 5% level whether the variance of machine A is greater than 0.04 mm2. (b) Test at the 5% level whether machine A's lengths vary more than machine B's.

0test statistic 20.0chi-sq(10) if H0 is trueH1: variance of A > 0.0410 × 0.08/0.04 = 20.0 on chi-sq(10)
For machine A the statistic (n − 1)s2σ2 = 20.0 follows χ2(10) if H0 is true.
For machine A, H0: σA2 = 0.04 and H1: σA2 > 0.04. The test statistic is (n − 1)sA20.04 = 10 × 0.080.04 = 20.0, on the χ2(10) distribution.
step 1 of 5

For a sample of n from a normal population with variance σ2, the quantity (n − 1)s2σ2 follows the χ2 distribution with n − 1 degrees of freedom, which gives a test of one variance. To compare two variances, the ratio of the two estimates follows the F distribution, with the degrees of freedom of the top estimate first and those of the bottom estimate second.

  1. For machine A, H0: σA2 = 0.04 and H1: σA2 > 0.04. The test statistic is (n − 1)sA20.04 = 10 × 0.080.04 = 20.0, on the χ2(10) distribution.
  2. (a) The upper 5% critical value of χ2(10) is 18.307. Since 20.0 > 18.307, reject H0: there is evidence at the 5% level that machine A's variance is greater than the 0.04 mm2 the specification allows.
  3. For the comparison, H0: σA2 = σB2 and H1: σA2 > σB2. The test statistic is the ratio of the estimates, the larger on top: F = 0.080.032 = 2.5.
  4. The degrees of freedom are 11 − 1 = 10 for the top and 13 − 1 = 12 for the bottom. The upper 5% critical value of F(10, 12) is 2.753.
  5. (b) 2.5 < 2.753, so do not reject H0. There is not enough evidence at the 5% level that machine A's lengths vary more than machine B's. The two results are consistent: A is shown to be outside the specification, but samples of 11 and 13 bolts are too small to show that A is worse than B.

answer(a) the test statistic is 20.0, above the critical value 18.307 of χ2(10), so reject H0: machine A's variance is greater than 0.04 mm2; (b) F = 2.5, below the critical value 2.753 of F(10, 12), so do not reject H0: there is not enough evidence that A's lengths vary more than B's

techniqueTesting a Variance with Chi-Squared · Comparing Two Variances with the F-Test

Common pitfalls

  • Using the standard deviations in place of the variances: 10 × √0.080.2 or F = √0.08√0.032 = 1.58. Both statistics are built from variances, s2 and σ2.
  • Swapping the degrees of freedom and reading F(12, 10), whose critical value is 2.913. The first number belongs to the estimate on top, machine A's, with 10 degrees of freedom.
Mr. Chalk Read the guide