Many lines, one best
Five students’ hours studied x and scores y are (1, 2), (3, 4), (5, 5), (7, 7) and (9, 8). The points rise from left to right, close to a straight line but not exactly on one.
Many straight lines pass near them. To use one for predicting a score from the hours, the best one has to be chosen by a rule, not by eye.
Residuals
For any line, a residual is the observed y minus the y the line gives at the same x. It is the vertical gap from the point to the line: positive for a point above the line, negative for a point below it.
The gaps are measured vertically because the line will be used to predict y from x, and a vertical gap is the error in y that such a prediction makes.
For the line y = 0.75x + 1.45, the line gives 2.2, 3.7, 5.2, 6.7 and 8.2 at x = 1, 3, 5, 7 and 9. The residuals are 2 − 2.2 = −0.2, 4 − 3.7 = 0.3, −0.2, 0.3 and −0.2.
The line y = 0.75x + 1.45 with the residuals drawn as vertical gaps: −0.2, 0.3, −0.2, 0.3 and −0.2. Their squares add up to 0.30.
Why the residuals are squared
Adding the residuals as they are does not pick out a line. The flat line y = 5.2, at the mean score, has residuals −3.2, −1.2, −0.2, 1.8 and 2.8. They add up to 0, just as −0.2 + 0.3 − 0.2 + 0.3 − 0.2 = 0 does, though the flat line misses every point by far more.
Squaring makes every gap count as positive, so gaps above and below the line cannot cancel. The flat line’s squared residuals add up to 10.24 + 1.44 + 0.04 + 3.24 + 7.84 = 22.8; those of y = 0.75x + 1.45 add up to 0.04 + 0.09 + 0.04 + 0.09 + 0.04 = 0.30.
Tilt the line to y = 0.4x + 3 and some gaps shrink while others grow. Its residuals are −1.4, −0.2, 0, 1.2 and 1.4, and their squares add up to 1.96 + 0.04 + 0 + 1.44 + 1.96 = 5.40.
The flatter line y = 0.4x + 3. Its residuals are −1.4, −0.2, 0, 1.2 and 1.4: long gaps at both ends, with squares adding up to 5.40.
The least-squares line
The least-squares regression line of y on x is the one line that makes the sum of the squared residuals as small as possible. For these five points it is y = 0.75x + 1.45, with total 0.30. Every other line gives more: y = 0.75x + 1.5 gives 0.3125, and y = 0.8x + 1.2 and y = 0.7x + 1.7 each give 0.40.
A calculator finds it from the list of pairs. By hand, written as y = ax + b, the gradient is a = Sxy/Sxx and the line passes through the mean point (x̄, ȳ), so b = ȳ − a x̄.
Here and ȳ . The x deviations are −4, −2, 0, 2 and 4, and the y deviations −3.2, −1.2, −0.2, 1.8 and 2.8. So Sxx = 16 + 4 + 0 + 4 + 16 = 40 and Sxy = 12.8 + 2.4 + 0 + 3.6 + 11.2 = 30.
and b = 5.2 − 0.75 × 5 = 5.2 − 3.75 = 1.45, so the line is y = 0.75x + 1.45. The residuals of the least-squares line always add up to 0, as these do.
the quantity being minimized is Σ(y − ŷ)², an area, so a distant point counts its distance squared: 4.73
Make the total area as small as you can
Six other points, from (1, 2) to (6, 6.5), with each residual’s square drawn on it. Drag the two ends of the line to make the total area as small as you can. The handles move in steps, and the least they reach is 0.375, at y = x + 0.75; the exact least-squares line, y = 0.96x + 0.9, gives 0.34.
The gradient and the intercept
a = 0.75 is the gradient, a rate of change: each extra hour of study goes with 0.75 more marks, on average, for this group.
b = 1.45 is the intercept, the score the line gives at x = 0, no study at all. The data run only from 1 hour to 9, so x = 0 lies outside them: b is a reading of the line beyond the points it was fitted to.
The least-squares line y = 0.75x + 1.45. It rises 0.75 for each hour and meets the score axis at 1.45, left of every point.
The usual mistakes
Adding the residuals without squaring. Their total is 0 for many different lines, the flat line at ȳ among them.
Measuring the gaps another way. The line of y on x minimizes the squared vertical gaps, not the horizontal gaps and not the shortest distances to the line.
Dividing by Syy instead of Sxx for the gradient. a = Sxy/Sxx.
Mixing up a and b. a is the change in y for each unit of x; b is the value of y when x is 0.
A print shop’s charges
In the application below, eight print jobs give a line from the sums , , and , and its gradient and intercept are read as the cost of each extra hundred copies and a fixed charge.
Worked example: A Print Shop's Charges for Eight Jobs, Where the Gradient Is the Cost of Each Extra Hundred Copies
Question A print shop charges $y for a job of x hundred copies. Eight recent jobs were: x = 2, 3, 5, 6, 8, 10, 12, 14 and y = 19.50, 26.50, 33.50, 40.00, 49.00, 57.00, 67.00, 73.50. You may use ∑ x = 60, ∑ y = 366, ∑ x2 = 578 and ∑ xy = 3321. (a) Find the equation of the least-squares regression line of y on x. (b) Explain what the gradient and the intercept mean for the shop's charges, and say which of the two is less reliable.
1.Work out the two sums the gradient needs: Sxx = 578 − 6028 = 578 − 450 = 128 and Sxy = 3321 − 60 × 3668 = 3321 − 2745 = 576.
The eight jobs, with the copies across and the charge up. Sxx = 128 and Sxy = 576. 2.The gradient is b = SxySxx = 576128 = 4.5.
The gradient is b = 576128 = 4.5. The line is drawn across the data only, from 2 to 14 hundred copies. 3.The line passes through the mean point (x, y) = (7.5, 45.75), so the intercept is a = y − bx = 45.75 − 4.5 × 7.5 = 45.75 − 33.75 = 12.
The line passes through the mean point, the cross at (7.5, 45.75), so a = 12. 4.(a) The regression line of y on x is y = 4.5x + 12.
(a) The regression line of y on x is y = 4.5x + 12. 5.(b) The gradient 4.5 means that each extra hundred copies adds about $4.50 to the charge, which is 4.5 cents a copy. The intercept 12 is the charge for no copies at all: a fixed charge of about $12 for setting up a job.
(b) The gradient triangle: 4 hundred more copies add $18, so each hundred adds $4.50. 6.The intercept is less reliable. The jobs start at 200 copies, so x = 0 lies outside the data and the fixed charge is read off the line beyond them. Check: at x = 8 the line gives 4.5 × 8 + 12 = $48, close to the $49 charged.
The intercept, dashed back to x = 0, is a fixed charge of $12, read beyond the smallest job.
Answer: (a) y = 4.5x + 12; (b) each extra hundred copies costs about $4.50, and there is a fixed charge of about $12 a job; the intercept is less reliable because x = 0 lies outside the data
Common mistakes
- Reading the gradient as the price of one copy, $4.50. x counts hundreds of copies, so $4.50 is the cost of a hundred more, which is 4.5 cents a copy.
- Dividing by Syy instead of Sxx for the gradient. The line of y on x measures its residuals in y, and its gradient is SxySxx.
More correlation and regression problems, worked step by step →