A straight line through the counts
A gardener counts the seedlings of a weed in a bed once a week. In weeks 1 to 5 the counts are 2, 4, 9, 15 and 33 plants. A model is a rule that predicts the count from the week, and the first one to try is a straight line.
The line of best fit through these counts is y = 7.3x − 9.3, where x is the week and y is the number of plants. It passes through the mean point, (3, 12.6), since 7.3 × 3 − 9.3 = 12.6. The question is whether a straight line is the right shape for this data at all.
Five weekly counts, with the best straight line, y = 7.3x − 9.3, laid over them.
Residuals: what the model missed
At each point, the residual is the actual value minus the value the model predicted:
residual = count − prediction
In week 1 the line predicts 7.3 × 1 − 9.3 = −2 plants, and the count was 2, so the residual is 2 − (−2) = +4. In week 4 the line predicts 7.3 × 4 − 9.3 = 19.9 plants, and the count was 15, so the residual is 15 − 19.9 = −4.9.
A positive residual means the point is above the model, which predicted too little. A negative residual means the point is below the model, which predicted too much. On the graph, each residual is the vertical gap from the point to the line.
The gap from each point straight up or down to the line is its residual. The line itself stops where it reaches 0 plants, just after week 1, so the gap at week 1 is not drawn.
A pattern in the residuals
Work out the residual for every week. The line’s predictions for weeks 1 to 5 are −2, 5.3, 12.6, 19.9 and 27.2, so the residuals are +4, −1.3, −3.6, −4.9 and +5.8. They add up to 0, as the residuals from a line through the mean point always do.
Read their signs in order: high, low, low, low, high. The line is below the points at both ends and above them in the middle. That is not random scatter. It is the shape of a curve bending upward, which a straight line cannot follow.
A model that fits leaves residuals that are small and have no pattern: their signs change without any order. Residuals that run in a pattern show that the model has missed part of the shape.
The line also fails a simpler test. It predicts −2 plants in week 1, and a count of plants cannot be negative.
The line’s residuals run +, −, −, −, +: the points are above the line at both ends and below it in the middle.
A doubling model
The counts roughly double each week: 2, 4, 9, 15, 33. A model that doubles each week is , which predicts 2, 4, 8, 16 and 32 plants for weeks 1 to 5. Its graph is a curve that gets steeper as it rises, the same shape the residuals of the line pointed to.
The doubling model, , follows the same five counts closely.
The residuals of the curve
The residuals of the curve are 2 − 2 = 0, 4 − 4 = 0, 9 − 8 = +1, 15 − 16 = −1 and 33 − 32 = +1. None is bigger than 1, and their signs change with no pattern, so there is no shape left over for another model to explain.
The gaps from the points to the curve are at most 1 plant, and they point up and down with no pattern.
The curve’s residuals are 0, 0, +1, −1 and +1: small, and with no pattern in their signs.
Comparing the sizes: the sum of squared residuals
To compare two models by one number, the residuals cannot simply be added, because the positive and negative ones cancel: the line’s residuals add up to 0, however badly it follows the shape. Square each residual first, as for the standard deviation, since a square is never negative, and then add the squares. The model with the smaller sum of squared residuals passes closer to the points.
For the line: . For the curve: . The line’s total is about 29 times the curve’s.
Does the model make sense?
The closer fit is only half of the choice. A model must also give sensible values for the situation. The line gives a negative number of plants whenever 7.3x − 9.3 < 0, that is, whenever , which includes week 1. The curve never goes below 0, and at week 0 it gives plant, a sensible start.
So the doubling model is the better choice, on both counts: its residuals are small with no pattern, and its values make sense. When two models fit about equally well, the simpler one, or the one that makes sense in the context, is chosen.
No model fits forever. Doubling every week would give plants in week 10 and over a million in week 20, and a flower bed has room for far fewer. A real population levels off as space runs out, so far from the data a model that levels off would be needed. Using a model beyond the data is the subject of the next lesson, Interpolation and Extrapolation.
Worked example: Braking Distance Against Speed, with a Straight Line and a Curve Both Offered
Question Six braking tests were made on one car. The pairs of (speed in km/h, braking distance in meters) were (20, 5), (30, 9), (40, 15), (50, 26), (60, 35) and (70, 48). Two models are offered: model A is y = 0.9x − 17 and model B is y = x2100. (a) Find the sum of the squared residuals for each model. (b) Say which model should be used, and test it against what the models give at rest.
1.Work out model A at each speed: 0.9 × 20 − 17 = 1, then 10, 19, 28, 37 and 46 meters.
The six tests, with model A drawn straight through them. The faint line across is zero meters. 2.Take the residuals for model A: 5 − 1 = 4, 9 − 10 = −1, 15 − 19 = −4, 26 − 28 = −2, 35 − 37 = −2 and 48 − 46 = 2. Squaring them gives 16 + 1 + 16 + 4 + 4 + 4 = 45.
Each dashed bar is a residual: the measured distance minus what model A gives at that speed. Squared, they total 45. 3.Work out model B at each speed: 202100 = 4, then 9, 16, 25, 36 and 49 meters. Its residuals are 1, 0, −1, 1, −1 and −1, and the squares total 1 + 0 + 1 + 1 + 1 + 1 = 5.
The same six residuals for model B, the curve y = x2100. Squared, they total 5. 4.(a) The sum of the squared residuals is 45 for model A and 5 for model B. Read the signs as well: model A is below the points at both ends and above them in the middle, and residuals that change sign in that pattern are the mark of a curve being fitted by a straight line.
(a) Model A totals 45 and model B totals 5. Model A also passes under the points at both ends and over them in the middle. 5.(b) Use model B. At rest the car needs no braking distance at all: model B gives 0100 = 0 meters, while model A gives 0.9 × 0 − 17 = −17 meters, and a distance cannot be negative. Check: model B says that doubling the speed from 30 to 60 km/h multiplies the braking distance by 4, from 9 to 36 meters, which is what braking distance is known to do.
(b) At rest, marked on the distance axis, model A gives −17 meters and model B gives 0 meters, so model B is the one to use.
Answer: (a) the sum of the squared residuals is 45 for model A and 5 for model B; (b) model B, y = x2100, because model A gives a braking distance of −17 meters at rest
Common mistakes
- Adding the residuals without squaring them. The positive and the negative residuals then cancel, and a line drawn through the middle of the points totals nearly zero however badly it follows their shape.
- Choosing model A because a straight line is the simpler rule. Simplicity decides only between models that fit equally well, and this one misses by 4 meters at the slowest test and gives a negative distance at rest.
More scatter plots and correlation problems, worked step by step →