A straight line of the wrong shape
Five plots are given 1, 3, 5, 7 and 9 units of water, and the crops are 3.8, 6.2, 7, 6.2 and 3.8. Too little water and too much both give a small crop, so the points rise to a peak and fall back.
The least-squares line through them is level, y = 5.4, the mean crop. It lies above the two outer points and below the middle three, and its residuals are −1.6, 0.8, 1.6, 0.8 and −1.6. Their squares total 2.56 + 0.64 + 2.56 + 0.64 + 2.56 = 8.96. The line does no better than the mean, so .
The best straight line, y = 5.4. The two outer points lie below it and the middle three above it, a pattern no straight line can remove.
The same rule, over a curve
Least squares does not need a straight line. For a quadratic model , it chooses a, b and c to make the total of the squared vertical gaps as small as possible, exactly as it chose the gradient and intercept of a line.
For these points the result is , which is . At x = 1 it gives −0.2 + 2 + 2 = 3.8, at x = 3 it gives −1.8 + 6 + 2 = 6.2, and at x = 5 it gives 7. These points lie exactly on the parabola, so every residual is 0 and the squared gaps total 0, against 8.96 for the line.
The least-squares quadratic, , passes through all five points, so no gap is left.
Five families of curves
A calculator offers five families besides the line, each with its own shape. A quadratic, , has one turning point. A cubic, cx + d, can turn twice. An exponential, , multiplies by the same factor b at every step of 1 in x, so it climbs or falls ever faster and never turns back. A power model, , such as or , passes through the origin when b is positive. A sine model, a sin(bx + c) + d, rises and falls in the same cycle again and again.
The data’s shape, and what is known about the situation, decide the family. Least squares then decides the numbers in it. Monthly rainfall that rises and falls in the same cycle every year calls for a sine model; a cubic turns at most twice and could not repeat the second year.
Counts that double
A count of colonies is 0.5 thousand at week 0, then 1, 2, 4 and 8 thousand at weeks 2, 4, 6 and 8. It doubles every two weeks, so it is multiplied by the same factor over every equal step: the exponential family, .
Here . Since , that is with a = 0.5 and , about 1.414: the count grows by 41.4 percent each week.
The curve through the five counts. Each step of two weeks doubles the count, so the curve gets steeper and steeper.
Straightening an exponential with logarithms
Take logarithms to base 10 of both sides of : log y = log a + x log b. That is a straight line, with log y up the side and x, unchanged, across. Its gradient is log b and its intercept is log a. Only y is logged.
For the colonies, log y at weeks 0, 2, 4, 6 and 8 is −0.301, 0, 0.301, 0.602 and 0.903. It rises by 0.301 every two weeks, so the points lie on a line with gradient 0.301 ÷ 2 = 0.1505 and intercept −0.301.
Read the model back: log b = 0.1505, so , and log a = −0.301, so . The gradient is not b and the intercept is not a; each is the logarithm of one.
The same five counts with log y plotted against the week. The points lie on a straight line with gradient 0.1505 and intercept −0.301.
A power model needs both logged
For a power model , logarithms give log y = log a + n log x. Now the line has log x across and log y up the side, so both variables are logged, and the gradient is n itself.
Take x = 1, 2, 4, 8 and y = 3, 12, 48, 192. The values of log x are 0, 0.301, 0.602 and 0.903, and of log y 0.477, 1.079, 1.681 and 2.283. The gradient is (2.283 − 0.477) ÷ 0.903 = 2, so n = 2, and the intercept is 0.477, so . The model is .
A line fitted to the logged values makes the squared gaps in log y smallest, not the squared gaps in y. The two fits are usually close, and they agree exactly when the points lie on the curve, as these do.
Transform back before predicting
To predict the colonies at week 10, use the straight line first: log y = −0.301 + 0.1505 × 10 = 1.204. That is the logarithm of the count, not the count. Then thousand, which is 8 thousand doubled, as two more weeks should give.
Week 10 lies beyond the data, which run to week 8, so 16 thousand assumes the doubling carries on.
The rule never changes
For the arch, the line leaves squared gaps totaling 8.96 and the quadratic leaves 0. Least squares chose the best member of each family; neither number says which family to choose. That choice comes from the shape of the data and from what the quantities are.
A family that contains another, as the quadratics contain every straight line (a = 0), always leaves smaller or equal squared gaps on the same data, so a smaller total is not enough on its own to prefer it. A quadratic is right for the arch because the crop rises to a peak and falls, not only because it fits better.
The usual mistakes
Fitting a straight line to data that curve. The residuals then follow a pattern, here negative at both ends and positive in the middle.
Reading the gradient of log y against x as b. A gradient of 0.1505 means .
Logging x as well as y for an exponential model, or only y for a power model. needs log y against x; needs log y against log x.
Giving the value on the straight line as the prediction. log y = 1.204 is a logarithm; the count is .
Bacteria counted every hour
In the application below, a count N of bacteria is modeled by . Logarithms to base 10 turn it into the line log N = log a + x log b, the line is fitted by least squares, and a and b are read back as powers of 10.
Worked example: Bacteria Counted Every Hour in a Culture, Fitted by N = ab^x Through Logarithms
Question A biologist counts the bacteria in a drop of culture every hour. At x = 0, 1, 2, 3, 4 and 5 hours the counts N are 129, 245, 501, 1000, 1950 and 4074. She expects a model of the form N = abx. (a) Using logarithms to base 10, find the least-squares regression line of log N on x. (b) Find a and b, and estimate the count after 4.5 hours.
1.Take logarithms to base 10 of both sides: log N = log a + x log b. This is a straight line in x and log N, with gradient log b and intercept log a.
The counts against the hours curve upward, so a straight line does not fit them. 2.To 2 decimal places, the values of log N are 2.11, 2.39, 2.70, 3.00, 3.29 and 3.61. They rise by about 0.3 each hour, so the transformed points lie close to a straight line.
The logarithms of the counts against the hours: the curve has become a straight line. 3.Write Y = log N. Then ∑ x = 15, ∑ Y = 17.1, ∑ x2 = 55 and ∑ xY = 48, so Sxx = 55 − 1526 = 17.5 and SxY = 48 − 15 × 17.16 = 48 − 42.75 = 5.25. The gradient is 5.2517.5 = 0.3.
The cross is the mean point (2.5, 2.85), and the gradient is 5.2517.5 = 0.3. 4.(a) The intercept is Y − 0.3x = 2.85 − 0.3 × 2.5 = 2.1, so the line is log N = 2.1 + 0.3x.
(a) The line through the mean point is log N = 2.1 + 0.3x. 5.(b) log a = 2.1 gives a = 102.1 = 126 to 3 significant figures, and log b = 0.3 gives b = 100.3 = 2.00. The model is N = 126 × 2.00x, so the count doubles every hour.
(b) a = 102.1 = 126 and b = 100.3 = 2.00: the count doubles every hour. 6.At x = 4.5: log N = 2.1 + 0.3 × 4.5 = 3.45, so N = 103.45 = 2820 to 3 significant figures. Check: at x = 3 the line gives log N = 3, which is N = 1000, the count recorded.
Reading up at 4.5 hours gives log N = 3.45, so N = 2820.
Answer: (a) log N = 2.1 + 0.3x; (b) a = 126 and b = 2.00, so N = 126 × 2.00x; after 4.5 hours about 2820 bacteria
Common mistakes
- Fitting a straight line to N itself. The counts double every hour, so they curve upward, and a straight line would lie above the middle counts and below the last one.
- Reading b as the gradient 0.3 and a as the intercept 2.1. The gradient is log b and the intercept is log a, so b = 100.3 = 2.00 and a = 102.1 = 126.
More correlation and regression problems, worked step by step →