The Coefficient of Determination

The share of the variation a model accounts for.

The variation with no model

Five students studied for 1, 3, 5, 7 and 9 hours and scored 3, 3, 6, 5 and 8. Without using the hours at all, the best single guess for a score is the mean score, 5.

The gaps from the mean are −2, −2, 1, 0 and 3. Their squares total 4 + 4 + 1 + 0 + 9 = 18. This total, Σ(y − ȳ)², is the variation in y: how far the scores spread around their mean.

hours studiedscore

The level line is the mean score, 5. The gaps from the points to it, squared and added, total 18.

The variation the line leaves

The least-squares line is y = 0.6x + 2. At 1, 3, 5, 7 and 9 hours it predicts 2.6, 3.8, 5, 6.2 and 7.4, so the residuals are 0.4, −0.8, 1, −1.2 and 0.6.

Their squares total 0.16 + 0.64 + 1 + 1.44 + 0.36 = 3.6. Using the hours has cut the squared gaps from 18 to 3.6.

hours studiedscore

The same points with the regression line y = 0.6x + 2. The gaps are shorter, and their squares total 3.6.

The share explained

Of the 18, the line leaves 3.6 unexplained, a share of 3.6 ÷ 18 = 0.2. The rest, 18 − 3.6 = 14.4, is accounted for by the line.

The coefficient of determination is R² = 1 − (squared residuals ÷ squared gaps from the mean) = 1 − 3.6 ÷ 18 = 0.8. The line accounts for 80 percent of the variation in the scores, and 20 percent is left in the residuals.

R² runs from 0 to 1. It is 0 when the model does no better than the mean, and 1 when every point lies on the model.

R² and r

For a straight-line fit, R² is exactly the square of Pearson’s r. Here Sxx = 40, Syy = 18 and Sxy = 24, so r = 24 ÷ √(40 × 18) = 24 ÷ √720 = 0.894, and r² = 576 / 720 = 0.8.

Square r before rounding it. Rounded first, 0.894² = 0.799, which is not quite 0.8.

The two numbers answer different questions. r has a sign and gives the direction and strength of the straight-line pattern. R² has no sign, because a square cannot be negative, and gives a share of the variation. r = −0.9 and r = 0.9 both give R² = 0.81.

Because squaring a number between 0 and 1 makes it smaller, r looks stronger than the share it explains. r = 0.7 sounds high, but r² = 0.49: the line accounts for less than half of the variation.

A bigger model always scores higher

Least squares can fit a curve to the same five points, and an extra term can never raise the squared residuals, because the new model can always set that term to 0 and do as well as before. So R² never falls as terms are added.

For these points the line gives R² = 0.8, a quadratic 0.816, and a cubic 0.821. A quartic has five coefficients, enough to pass through all five points, so its squared residuals total 0 and R² = 1.

The quartic is not the better model. Between and beyond the points it swings to follow every one of them: at 0 hours it predicts a score of 10.4, higher than any student scored. A larger R² from extra terms is not evidence that the extra terms are real.

hours studiedscore

The quartic through all five points. Every residual is 0, so R² = 1, but the curve dips and rises between the points and starts at 10.4 at 0 hours.

Look at the residuals too

Another five students scored 2, 5, 7, 8 and 8.5 after 1, 3, 5, 7 and 9 hours. The line y = 0.8x + 2.1 gives R² = 0.91, which sounds like a good fit.

In order, its residuals are −0.9, 0.5, 0.9, 0.3 and −0.8: negative at both ends and positive in the middle. The points bend over, and the line cuts across the bend. A high R² says the line follows the points closely; the pattern in the residuals says a curve would follow them better.

hours studiedscore

R² = 0.91, but the first and last points lie below the line and the middle three above it. The gaps follow a curve.

The usual mistakes

Giving the unexplained share. 3.6 ÷ 18 = 0.2 is what the line misses; R² is 1 − 0.2 = 0.8.

Dividing the totals the wrong way up. 18 ÷ 3.6 = 5, and a share of the variation cannot be more than 1.

Reading R² as a share of the points. R² = 0.64 means 64 percent of the variation in y is accounted for, not that 64 percent of the points lie on the line.

Using r as R². For r = 0.8, R² = 0.64, not 0.8; for r = −0.6, R² = 0.36, not −0.36.

Flats sold in one street

In the application below, the price of ten flats is set against their floor area, with Syy = 90000 as the total variation in price. The line accounts for Sxy² ÷ Sxx of it, and the residuals hold the rest.

Worked example: Flats Sold in One Street: How Much of the Price the Floor Area Explains

Question An estate agent records the floor area x m2 and the selling price y thousand dollars of 10 flats sold in one street. The results give Sxx = 2500, Syy = 90000 and Sxy = 10500. The agent says that floor area decides the price. (a) Calculate r and the coefficient of determination r2. (b) Find the sum of the squared residuals about the regression line of y on x, and say what share of the variation in price the floor area explains.

  1. 1.Pearson's correlation coefficient is r = Sxy√Sxx Syy = 10500√2500 × 90000 = 1050015000 = 0.7.

    Syytotal variation 90000r = 10500/√(2500 × 90000)= 10500/15000 = 0.7
    Syy90000r = 10500/√(2500 × 90000)= 10500/15000 = 0.7
    The bar is the total variation in price, Syy = 90000. r = 1050015000 = 0.7.
  2. 2.(a) r = 0.7, and the coefficient of determination is r2 = 0.72 = 0.49.

    Syytotal variation 90000split49%?r2= 0.7 × 0.7 = 0.49
    Syy90000split49%?r2= 0.7 × 0.7 = 0.49
    (a) The coefficient of determination is r2 = 0.49.
  3. 3.The total variation in price is Syy = 90000. The regression line accounts for Sxy2Sxx = 1050022500 = 44100 of it.

    Syytotal variation 90000splitexplained 44100residuals ?explained: 10500 × 10500/2500 = 44100
    Syy90000split44100?explained: 10500 × 10500/2500 = 44100
    The line accounts for Sxy2Sxx = 44100 of the 90000.
  4. 4.(b) The squared residuals add up to the rest: 90000 − 44100 = 45900. Floor area explains 4410090000 = 0.49, that is 49% of the variation in price, and 51% is left to other things, such as the floor the flat is on and its condition.

    Syytotal variation 90000splitexplained 44100residuals 4590049%51%residuals: 90000 − 44100 = 4590049% explained, 51% not
    Syy90000split441004590049%51%residuals: 90000 − 44100 = 4590049% explained, 51% not
    (b) The residuals keep 45900: floor area explains 49% of the variation in price.
  5. 5.The agent's claim is too strong: r = 0.7 is a fairly strong correlation, but floor area accounts for less than half of the variation in price. Check: 4590090000 = 0.51 = 1 − r2.

    Syytotal variation 90000splitexplained 44100residuals 4590049%51%r = 0.7, but less than half explained
    Syy90000split441004590049%51%r = 0.7, but less than half explained
    A correlation of 0.7 leaves 51% of the variation to other things, so floor area does not decide the price.

Answer: (a) r = 0.7 and r2 = 0.49; (b) the squared residuals add up to 45900; floor area explains 49% of the variation in price, less than half

Common mistakes

  • Reading r = 0.7 as 70% of the variation explained. The share explained is r2 = 0.49, which is less than half.
  • Taking r2 = 0.49 to mean that the line is wrong for about half of the flats. It is a share of the variation, the total squared distance of the prices from their mean, not a count of flats.

More correlation and regression problems, worked step by step →

Practice The Coefficient of Determination in the app