A line fitted to a range
Five students’ hours of study and scores on a quiz marked out of 10 are (2, 3), (4, 4), (5, 6), (7, 6) and (8, 8). The hours run from 2 to 8.
The means are and ȳ . The x deviations are −3.2, −1.2, −0.2, 1.8 and 2.8, so Sxx = 10.24 + 1.44 + 0.04 + 3.24 + 7.84 = 22.8. The y deviations are −2.4, −1.4, 0.6, 0.6 and 2.6, so Sxy = 7.68 + 1.68 − 0.12 + 1.08 + 7.28 = 17.6.
The least-squares line of y on x has gradient and passes through (5.2, 5.4), so b = 5.4 − 0.772 × 5.2 = 1.386. The line is y = 0.772x + 1.386, and r = 0.945.
The line was chosen to fit these five points, all between 2 and 8 hours. What it says outside that range has not been tested against anything.
The five points, from 2 hours to 8, and the least-squares line y = 0.772x + 1.386. At 6 hours the line gives 6.02, between the points at 5 and 7 hours.
Inside the data: interpolation
To predict the score for 6 hours of study, put x = 6 into the line: y = 0.772 × 6 + 1.386 = 6.02, a score of about 6.
6 hours lies inside the range of the data, between the points at 5 and 7 hours, which both score 6. Reading the line between measured points is interpolation. The points close by show how well the line fits there, so with r = 0.945 the prediction is reliable.
Outside the data: extrapolation
At 18 hours the same line gives y = 0.772 × 18 + 1.386 = 15.28, a score of more than 15 on a quiz marked out of 10. The arithmetic is right; the prediction is impossible.
18 hours lies far beyond the data. Reading the line outside the range is extrapolation. It assumes the pattern seen between 2 and 8 hours carries on, and nothing in the data shows that. Here it cannot: scores stop at 10.
A strong correlation does not help. r = 0.945 describes how closely the five points follow the line between 2 and 8 hours, and says nothing about hours that were never measured.
The same line stretched to 20 hours. At 18 hours it gives 15.28, far above the 10 marks the quiz has, and every point is back between 2 and 8 hours.
A result and a guess
Inside the range of the data, a prediction is a result: the fit was measured there. Beyond it, at either end, a prediction is a guess: nothing shows the pattern continues.
The edges are not sharp. At 10 hours the line gives 0.772 × 10 + 1.386 = 9.11, just beyond the data and still a possible score, but no less an extrapolation. At 0 hours it gives the intercept, 1.386, and that is an extrapolation too, since the data start at 2 hours.
Which way the line predicts
The line of y on x was chosen to make the vertical gaps small, the errors in y. It predicts y from a given x.
Rearranging y = 0.772x + 1.386 to find the hours that go with a score of 7 uses the same line for a job it was not fitted for. Predicting x from y needs a different line, fitted to make the horizontal gaps small, which is a later lesson.
The usual mistakes
Trusting an extrapolation because the points lie close to the line. The fit is only known where there are points.
Rejecting an interpolation because no student studied exactly 6 hours. Between measured points the line is supported by the data on both sides.
Dropping the intercept. 0.772 × 6 = 4.63 is not the prediction; the line does not pass through the origin, so 1.386 is added.
Rearranging the line to predict x from y. The line of y on x predicts y only.
A girl’s height on each birthday
In the application below, heights measured from age 2 to age 10 give a line that is used at age 7.5, inside the data, and at age 30, far beyond it.
Worked example: A Girl's Height on Each Birthday from Two to Ten, Used Once Inside the Record and Once Twenty Years Beyond It
Question A girl's height y cm was measured on each birthday from age x = 2 to age x = 10 years. The nine heights were 86, 92, 102, 108, 113, 119, 125, 129 and 134 cm. The regression line of y on x is y = 6x + 76. (a) Use the line to estimate her height at age 7.5 years, and say whether the estimate is reliable. (b) Find the height the line gives at age 30, and explain why it should not be trusted.
1.The measurements run from age 2 to age 10, so the line is supported by the data only for ages between 2 and 10.
The nine birthdays, with the line drawn across the data only, from age 2 to age 10. 2.At x = 7.5: y = 6 × 7.5 + 76 = 45 + 76 = 121 cm.
Reading up at 7.5 years: 6 × 7.5 + 76 = 121 cm. 3.(a) Her estimated height at 7.5 years is 121 cm. The estimate is reliable: 7.5 lies inside the data, so this is interpolation, and the heights lie close to the line.
(a) 121 cm, interpolation inside the data, so the estimate is reliable. 4.At x = 30: y = 6 × 30 + 76 = 180 + 76 = 256 cm.
Carried on, dashed, to age 30, the line gives 6 × 30 + 76 = 256 cm. 5.(b) The line gives 256 cm, which is over 2.5 m and far taller than almost any adult. Age 30 is far outside the data, so this is extrapolation: the line assumes she grows 6 cm a year forever, but people stop growing in their late teens. Check: at the mean age 6 the line gives 6 × 6 + 76 = 112 cm, the mean of the nine heights.
(b) 256 cm is extrapolation twenty years beyond the data, and no one grows 6 cm a year forever.
Answer: (a) 121 cm, reliable because 7.5 lies inside the data; (b) 256 cm, which cannot be trusted because age 30 is far outside the data and people stop growing in their late teens
Common mistakes
- Trusting the prediction at age 30 because the points lie so close to the line. A strong correlation shows that the line fits ages 2 to 10; it says nothing about ages the data do not cover.
- Rejecting the estimate at 7.5 years because no measurement was taken at that age. Between measured ages the line is interpolation, which the data support.
More correlation and regression problems, worked step by step →