The Least-Squares Regression Line

The line that makes the squared gaps smallest.

Many lines, one best

Five students’ hours studied x and scores y are (1, 2), (3, 4), (5, 5), (7, 7) and (9, 8). The points rise from left to right, close to a straight line but not exactly on one.

Many straight lines pass near them. To use one for predicting a score from the hours, the best one has to be chosen by a rule, not by eye.

Residuals

For any line, a residual is the observed y minus the y the line gives at the same x. It is the vertical gap from the point to the line: positive for a point above the line, negative for a point below it.

The gaps are measured vertically because the line will be used to predict y from x, and a vertical gap is the error in y that such a prediction makes.

For the line y = 0.75x + 1.45, the line gives 2.2, 3.7, 5.2, 6.7 and 8.2 at x = 1, 3, 5, 7 and 9. The residuals are 2 − 2.2 = −0.2, 4 − 3.7 = 0.3, −0.2, 0.3 and −0.2.

hours studiedscore

The line y = 0.75x + 1.45 with the residuals drawn as vertical gaps: −0.2, 0.3, −0.2, 0.3 and −0.2. Their squares add up to 0.30.

Why the residuals are squared

Adding the residuals as they are does not pick out a line. The flat line y = 5.2, at the mean score, has residuals −3.2, −1.2, −0.2, 1.8 and 2.8. They add up to 0, just as −0.2 + 0.3 − 0.2 + 0.3 − 0.2 = 0 does, though the flat line misses every point by far more.

Squaring makes every gap count as positive, so gaps above and below the line cannot cancel. The flat line’s squared residuals add up to 10.24 + 1.44 + 0.04 + 3.24 + 7.84 = 22.8; those of y = 0.75x + 1.45 add up to 0.04 + 0.09 + 0.04 + 0.09 + 0.04 = 0.30.

Tilt the line to y = 0.4x + 3 and some gaps shrink while others grow. Its residuals are −1.4, −0.2, 0, 1.2 and 1.4, and their squares add up to 1.96 + 0.04 + 0 + 1.44 + 1.96 = 5.40.

hours studiedscore

The flatter line y = 0.4x + 3. Its residuals are −1.4, −0.2, 0, 1.2 and 1.4: long gaps at both ends, with squares adding up to 5.40.

The least-squares line

The least-squares regression line of y on x is the one line that makes the sum of the squared residuals as small as possible. For these five points it is y = 0.75x + 1.45, with total 0.30. Every other line gives more: y = 0.75x + 1.5 gives 0.3125, and y = 0.8x + 1.2 and y = 0.7x + 1.7 each give 0.40.

A calculator finds it from the list of pairs. By hand, written as y = ax + b, the gradient is a = Sxy/Sxx and the line passes through the mean point (x̄, ȳ), so b = ȳ − a x̄.

Here x̄ = 25/5 = 5 and ȳ = 26/5 = 5.2. The x deviations are −4, −2, 0, 2 and 4, and the y deviations −3.2, −1.2, −0.2, 1.8 and 2.8. So Sxx = 16 + 4 + 0 + 4 + 16 = 40 and Sxy = 12.8 + 2.4 + 0 + 3.6 + 11.2 = 30.

a = 30/40 = 0.75 and b = 5.2 − 0.75 × 5 = 5.2 − 3.75 = 1.45, so the line is y = 0.75x + 1.45. The residuals of the least-squares line always add up to 0, as these do.

total area 4.73

the quantity being minimized is Σ(y − ŷ)², an area, so a distant point counts its distance squared: 4.73

Make the total area as small as you can

Six other points, from (1, 2) to (6, 6.5), with each residual’s square drawn on it. Drag the two ends of the line to make the total area as small as you can. The handles move in steps, and the least they reach is 0.375, at y = x + 0.75; the exact least-squares line, y = 0.96x + 0.9, gives 0.34.

The gradient and the intercept

a = 0.75 is the gradient, a rate of change: each extra hour of study goes with 0.75 more marks, on average, for this group.

b = 1.45 is the intercept, the score the line gives at x = 0, no study at all. The data run only from 1 hour to 9, so x = 0 lies outside them: b is a reading of the line beyond the points it was fitted to.

hours studiedscore

The least-squares line y = 0.75x + 1.45. It rises 0.75 for each hour and meets the score axis at 1.45, left of every point.

The usual mistakes

Adding the residuals without squaring. Their total is 0 for many different lines, the flat line at ȳ among them.

Measuring the gaps another way. The line of y on x minimizes the squared vertical gaps, not the horizontal gaps and not the shortest distances to the line.

Dividing by Syy instead of Sxx for the gradient. a = Sxy/Sxx.

Mixing up a and b. a is the change in y for each unit of x; b is the value of y when x is 0.

A print shop’s charges

In the application below, eight print jobs give a line from the sums Σx, Σy, Σx² and Σxy, and its gradient and intercept are read as the cost of each extra hundred copies and a fixed charge.

Worked example: A Print Shop's Charges for Eight Jobs, Where the Gradient Is the Cost of Each Extra Hundred Copies

Question A print shop charges $y for a job of x hundred copies. Eight recent jobs were: x = 2, 3, 5, 6, 8, 10, 12, 14 and y = 19.50, 26.50, 33.50, 40.00, 49.00, 57.00, 67.00, 73.50. You may use ∑ x = 60, ∑ y = 366, ∑ x2 = 578 and ∑ xy = 3321. (a) Find the equation of the least-squares regression line of y on x. (b) Explain what the gradient and the intercept mean for the shop's charges, and say which of the two is less reliable.

  1. 1.Work out the two sums the gradient needs: Sxx = 578 − 6028 = 578 − 450 = 128 and Sxy = 3321 − 60 × 3668 = 3321 − 2745 = 576.

    0204060800481216charge, $copies, hundredsSxx = 578 − 450 = 128Sxy = 3321 − 2745 = 576
    0204060800481216charge, $copies, hundredsSxx = 578 − 450 = 128Sxy = 3321 − 2745 = 576
    The eight jobs, with the copies across and the charge up. Sxx = 128 and Sxy = 576.
  2. 2.The gradient is b = SxySxx = 576128 = 4.5.

    0204060800481216charge, $copies, hundredsb = 576/128 = 4.5
    0204060800481216charge, $copies, hundredsb = 576/128 = 4.5
    The gradient is b = 576128 = 4.5. The line is drawn across the data only, from 2 to 14 hundred copies.
  3. 3.The line passes through the mean point (x, y) = (7.5, 45.75), so the intercept is a = y − bx = 45.75 − 4.5 × 7.5 = 45.75 − 33.75 = 12.

    0204060800481216charge, $copies, hundredsmean point (7.5, 45.75)a = 45.75 − 4.5 × 7.5 = 12
    0204060800481216charge, $copies, hundredsmean point (7.5, 45.75)a = 45.75 − 4.5 × 7.5 = 12
    The line passes through the mean point, the cross at (7.5, 45.75), so a = 12.
  4. 4.(a) The regression line of y on x is y = 4.5x + 12.

    0204060800481216charge, $copies, hundredsy = 4.5x + 12
    0204060800481216charge, $copies, hundredsy = 4.5x + 12
    (a) The regression line of y on x is y = 4.5x + 12.
  5. 5.(b) The gradient 4.5 means that each extra hundred copies adds about $4.50 to the charge, which is 4.5 cents a copy. The intercept 12 is the charge for no copies at all: a fixed charge of about $12 for setting up a job.

    0204060800481216charge, $copies, hundreds4 hundred$184 hundred more copies: $18 more$4.50 a hundred, 4.5 cents a copy
    0204060800481216charge, $copies, hundreds4 hundred$184 hundred more copies: $18 more$4.50 a hundred, 4.5 cents a copy
    (b) The gradient triangle: 4 hundred more copies add $18, so each hundred adds $4.50.
  6. 6.The intercept is less reliable. The jobs start at 200 copies, so x = 0 lies outside the data and the fixed charge is read off the line beyond them. Check: at x = 8 the line gives 4.5 × 8 + 12 = $48, close to the $49 charged.

    0204060800481216charge, $copies, hundreds4 hundred$18at x = 0 the line gives $12x = 0 is outside the data
    0204060800481216charge, $copies, hundreds4 hundred$18at x = 0 the line gives $12x = 0 is outside the data
    The intercept, dashed back to x = 0, is a fixed charge of $12, read beyond the smallest job.

Answer: (a) y = 4.5x + 12; (b) each extra hundred copies costs about $4.50, and there is a fixed charge of about $12 a job; the intercept is less reliable because x = 0 lies outside the data

Common mistakes

  • Reading the gradient as the price of one copy, $4.50. x counts hundreds of copies, so $4.50 is the cost of a hundred more, which is 4.5 cents a copy.
  • Dividing by Syy instead of Sxx for the gradient. The line of y on x measures its residuals in y, and its gradient is SxySxx.

More correlation and regression problems, worked step by step →

Practice The Least-Squares Regression Line in the app