The line of y on x
Five students record the hours x they studied and the score y they got: (1, 2), (3, 6), (5, 4), (7, 6) and (9, 7). The least-squares line of y on x is y = 0.5x + 2.5. It makes the total of the squared vertical gaps, the residuals in y, as small as possible.
Its residuals are 2 − 3 = −1, 6 − 4 = 2, 4 − 5 = −1, 6 − 6 = 0 and 7 − 7 = 0, and their squares total 1 + 4 + 1 + 0 + 0 = 6. The gaps are vertical because the line is built to predict y: for 7 hours of study it gives a score of 0.5 × 7 + 2.5 = 6.
The line of y on x, y = 0.5x + 2.5, and the vertical gaps from the points to it. Their squares total 6, the smallest total any straight line can give.
Predicting the hours instead
Now a student’s score is known and the hours are to be estimated. The error that matters is now an error in x, so the line should make the squared horizontal gaps as small as possible. That is the line of x on y, written x = cy + d.
It is found the same way, with the roles of x and y swapped. With Sxy ȳ), Sxx and Syy ȳ, the gradient of y on x is Sxy ÷ Sxx and the gradient of x on y is Sxy ÷ Syy. Both means are 5, and the sums are Sxx = 40, Syy = 16 and Sxy = 20.
So the gradient of y on x is 20 ÷ 40 = 0.5, and the gradient of x on y is . The line passes through the mean point, so 5 = 1.25 × 5 + d, which gives d = −1.25. The line of x on y is x = 1.25y − 1.25.
The same students with score across and hours up. The line is x = 1.25y − 1.25, and the gaps drawn to it are the horizontal gaps of the original plot, now standing upright. Their squares total 15.
Two different lines
Back on the original axes, with hours across and score up, x = 1.25y − 1.25 rearranges to . It is steeper than y = 0.5x + 2.5, so the two lines are different.
Each line wins only its own contest. The line of y on x leaves squared vertical gaps totaling 6, while the line of x on y leaves 9.6. The line of x on y leaves squared horizontal gaps totaling 15, while the line of y on x leaves 24.
Both lines from the same five points: y on x, y = 0.5x + 2.5, and x on y, which is y = 0.8x + 1 on these axes. They cross at the mean point (5, 5).
Using the line of x on y
A sixth student scored 6. Put y = 6 into x = 1.25y − 1.25: x = 1.25 × 6 − 1.25 = 7.5 − 1.25 = 6.25, so the estimate is 6.25 hours.
Rearranging the line of y on x gives a different answer. From 6 = 0.5x + 2.5, hours. That line was chosen to make the errors in the score small, not the errors in the hours, so it is the wrong line for this estimate.
Where the two lines meet
Both lines pass through the mean point (x̄, ȳ). Here y = 0.5 × 5 + 2.5 = 5 and x = 1.25 × 5 − 1.25 = 5, so both pass through (5, 5).
They are the same line only when r = 1 or r = −1. The product of the two gradients, (Sxy ÷ Sxx) × (Sxy ÷ Syy), is Sxy² ÷ (Sxx × Syy), which is . Here 0.5 × 1.25 = 0.625, so and r = 0.791.
On the same axes the line of x on y has gradient 1 ÷ 1.25 = 0.8. The two gradients 0.5 and 0.8 agree only when 0.5 × 1.25 = 1, that is when , with every point on one line. The weaker the correlation, the wider the lines open. When r = 0 the line of y on x is level, y = ȳ, and the line of x on y is upright, x = x̄.
Which line to use
Choose by what is being predicted, and use the line that minimizes the residuals in that variable. To estimate y from a known x, use the line of y on x. To estimate x from a known y, use the line of x on y.
The usual mistakes
Rearranging the line of y on x to estimate x. It gives 7 hours for a score of 6, where the line of x on y gives 6.25.
Dividing by Sxx for the gradient of x on y. 20 ÷ 40 = 0.5 is the gradient of y on x; for x on y the divisor is Syy, giving 20 ÷ 16 = 1.25.
Dropping the constant. x = 1.25 × 6 = 7.5 leaves out the −1.25, and the line does not pass through the origin.
Treating the two lines as one. They meet only at the mean point unless r is 1 or −1.
A French mark from a Spanish mark
In the application below, one student missed the French exam. Her French mark x is estimated from her Spanish mark y, so the line of x on y is used, and its gradient is Sxy ÷ Syy.
Worked example: A French Mark Estimated from a Spanish Mark, for a Student Who Missed the French Exam
Question Ten students sat a French exam and a Spanish exam. Their French marks x and Spanish marks y give ∑ x = 550, ∑ y = 500, ∑ x2 = 30750, ∑ y2 = 27000 and ∑ xy = 28300. An eleventh student scored 60 in Spanish but was ill on the day of the French exam. (a) Find the regression line of x on y, and use it to estimate her French mark. (b) Find the regression line of y on x, rearrange it to give x, and explain why its estimate of her French mark is different.
1.Work out the three sums: Sxx = 30750 − 550210 = 30750 − 30250 = 500, Syy = 27000 − 500210 = 27000 − 25000 = 2000 and Sxy = 28300 − 550 × 50010 = 28300 − 27500 = 800.
The mean point (55, 50) is the cross. Both regression lines pass through it. 2.The French mark is estimated from a known Spanish mark, so use the line of x on y, x = c + dy, with d = SxySyy = 8002000 = 0.4. It passes through the mean point (x, y) = (55, 50), so c = 55 − 0.4 × 50 = 35.
The line of x on y has gradient d = 8002000 = 0.4 when x is written in terms of y: x = 35 + 0.4y. 3.(a) The line of x on y is x = 35 + 0.4y. At y = 60: x = 35 + 0.4 × 60 = 35 + 24 = 59, so her French mark is estimated as 59.
(a) Reading across from a Spanish mark of 60 to the line of x on y and down gives a French mark of 59. 4.The line of y on x has gradient b = SxySxx = 800500 = 1.6 and intercept a = 50 − 1.6 × 55 = 50 − 88 = −38, so it is y = 1.6x − 38.
The line of y on x, y = 1.6x − 38, is less steep: it is a different line through the same mean point. 5.(b) Rearranged, x = y + 381.6, which at y = 60 gives 981.6 = 61.25. It is different because the line of y on x minimizes the errors in the Spanish mark, while the French mark is the one being estimated. The two lines cross at the mean point and are the same line only when r = ± 1; here r = 800√500 × 2000 = 8001000 = 0.8.
(b) Rearranged, it gives 61.25 at y = 60. It minimizes the errors in y, so 59 is the estimate to use.
Answer: (a) x = 35 + 0.4y, which estimates her French mark as 59; (b) y = 1.6x − 38, which rearranged gives 61.25; that line minimizes the errors in the Spanish mark, not in the French mark being estimated, so 59 is the estimate to use
Common mistakes
- Rearranging the line of y on x to estimate x. That line was chosen to make the errors in y small, so it is the wrong line for estimating x unless the correlation is perfect.
- Dividing by Sxx for the gradient of x on y. 800500 = 1.6 is the gradient of y on x; for x on y the denominator is Syy, the spread of the variable that is known.
More correlation and regression problems, worked step by step →