Mean of Grouped Data

Estimate using each interval midpoint.

Only the counts are left

A gardener measures 10 seedlings. She does not keep each height; she records only the class it falls in. Writing h for a height in centimeters, 4 seedlings are in the class 0 ≤ h < 10 and 6 are in the class 10 ≤ h < 20.

The mean is the total of the values divided by how many values there are. But the values are gone. Her record says that 4 seedlings are shorter than 10 cm, and nothing about how much shorter, so there are no heights to add.

01234560–1010–20

The histogram shows how many seedlings are in each class, 4 and 6, and not where each height lies inside its class.

One table, many possible data sets

Here are two different sets of heights that both give this table. In the first, every seedling is near the bottom of its class: 1, 1, 2 and 2 cm, then 10, 11, 11, 12, 12 and 13 cm. Their total is 6 + 69 = 75 cm, and their mean is 75 ÷ 10 = 7.5 cm.

In the second, every seedling is near the top of its class: 8, 9, 9 and 9 cm, then 17, 18, 18, 19, 19 and 19 cm. Their total is 35 + 110 = 145 cm, and their mean is 145 ÷ 10 = 14.5 cm.

Both sets have 4 heights in the first class and 6 in the second, so the table cannot tell them apart, yet their means differ by 7 cm. From grouped data alone the mean cannot be found exactly. It can only be estimated.

05101520

The first set: 4 dots below 10 and 6 dots from 10 up to 20, all low in their classes. The mean is 7.5 cm.

05101520

The second set has the same counts, 4 and 6, with every dot high in its class. The mean is 14.5 cm.

Midpoints stand in for the values

The usual value to stand in for every value in a class is its midpoint, the number halfway between its two ends. The midpoint of 0 ≤ h < 10 is (0 + 10)/2 = 5, and the midpoint of 10 ≤ h < 20 is (10 + 20)/2 = 15.

So treat every seedling in the first class as 5 cm tall and every seedling in the second class as 15 cm tall. When the real heights are spread through each class, some lie above the midpoint and some below, and those errors partly cancel when the heights are added. Then average as usual.

05101520midmid

The midpoints, 5 and 15, sit halfway along the classes 0 to 10 and 10 to 20.

The estimated mean

With every value replaced by its midpoint, the data is 5 four times and 15 six times. Multiply each midpoint by its frequency and add: 5 × 4 + 15 × 6 = 20 + 90 = 110 cm. Divide by the number of seedlings: 110 ÷ 10 = 11 cm.

So the estimated mean height is 11 cm. Always say that it is an estimate. The two sets of heights with this same table have means of 7.5 cm and 14.5 cm, and the estimate, 11 cm, lies between them.

The estimate, 11, is closer to 15 than to 5, because the second class holds more seedlings. Each midpoint is weighted by its frequency. Taking the mean of the two midpoints, (5 + 15)/2 = 10, treats the two classes as if they held equally many seedlings, and they do not.

515mean 11

Four seedlings at 5 cm and six at 15 cm balance at 11, pulled toward the class with more seedlings in it.

Classes with unequal ends

Classes do not always start at round numbers or have the same width. The midpoint is always found the same way: add the two ends of the class and halve. For times t in minutes, the class 20 ≤ t < 35 has midpoint (20 + 35)/2 = 27.5, and the class 35 ≤ t < 60 has midpoint (35 + 60)/2 = 47.5.

Do not take 5 more than the start of each class out of habit: 25 is the midpoint of a class from 20 to 30, not of one from 20 to 35. A class 15 wide has its midpoint 7.5 above its start.

The modal class

With grouped data there is no single most common value, because the values are gone. What the table does show is the class with the largest frequency, called the modal class. For the seedlings it is 10 ≤ h < 20, with 6 seedlings. The modal class is an interval of heights; the 6 is its frequency, not the answer.

Worked example: Journeys to Work Grouped into Classes: the Modal Class and an Estimated Mean

Question A company asked all 40 of its staff how long the journey to work takes. Writing the time as m minutes, the answers were 0 < m ≤ 10 for 4 staff, 10 < m ≤ 20 for 10 staff, 20 < m ≤ 30 for 14 staff, 30 < m ≤ 40 for 8 staff and 40 < m ≤ 50 for 4 staff. (a) Write down the modal class and estimate the mean journey time. (b) Explain why that mean is only an estimate, and find the smallest and the largest value the true mean could have.

  1. 1.Write down the middle of each class, halfway between its two ends: 5, 15, 25, 35 and 45 minutes.

    minutesmiddle xstaff ff x0 to 105410 to 20151020 to 30251430 to 4035840 to 50454totalthe middle of 20 to 30 is 25
    minutesmiddle xstaff ff x0 to 105410 to 20151020 to 30251430 to 4035840 to 50454totalthe middle of 20 to 30 is 25
    Each class is replaced by the time at its middle: 5, 15, 25, 35 and 45 minutes.
  2. 2.Multiply each middle by its frequency: 5 × 4 = 20, 15 × 10 = 150, 25 × 14 = 350, 35 × 8 = 280 and 45 × 4 = 180.

    minutesmiddle xstaff ff x0 to 10542010 to 20151015020 to 30251435030 to 4035828040 to 50454180total20, 150, 350, 280 and 180 minutes
    minutesmiddle xstaff ff x0 to 10542010 to 20151015020 to 30251435030 to 4035828040 to 50454180total20, 150, 350, 280 and 180 minutes
    Multiply each middle by its frequency: 5 × 4 = 20, 15 × 10 = 150, and so on.
  3. 3.(a) The largest frequency is 14, so the modal class is 20 < m ≤ 30 minutes. The totals are 40 staff and 20 + 150 + 350 + 280 + 180 = 980 minutes, so the estimated mean is 980 ÷ 40 = 24.5 minutes.

    minutesmiddle xstaff ff x0 to 10542010 to 20151015020 to 30251435030 to 4035828040 to 50454180total40980980 minutes over 40 staff980 divided by 40 = 24.5 minutes
    minutesmiddle xstaff ff x0 to 10542010 to 20151015020 to 30251435030 to 4035828040 to 50454180total40980980 minutes over 40 staff980 divided by 40 = 24.5 minutes
    (a) The largest frequency is 14, so the modal class is 20 < m ≤ 30 minutes, and the estimated mean is 980 ÷ 40 = 24.5 minutes.
  4. 4.The estimate is only an estimate because every time was replaced by the middle of its class. To see how far out it could be, give everyone the lowest time their class allows: 0 × 4 + 10 × 10 + 20 × 14 + 30 × 8 + 40 × 4 = 780, and 780 ÷ 40 = 19.5 minutes.

    minutesmiddle xstaff ff x0 to 10542010 to 20151015020 to 30251435030 to 4035828040 to 50454180total40980give everyone the bottom of the class780 divided by 40 = 19.5 minutes
    minutesmiddle xstaff ff x0 to 10542010 to 20151015020 to 30251435030 to 4035828040 to 50454180total40980give everyone the bottom of the class780 divided by 40 = 19.5 minutes
    Give everyone the lowest time their class allows and the mean is 780 ÷ 40 = 19.5 minutes.
  5. 5.(b) Now give everyone the highest time their class allows: 10 × 4 + 20 × 10 + 30 × 14 + 40 × 8 + 50 × 4 = 1180, and 1180 ÷ 40 = 29.5 minutes. The true mean therefore lies between 19.5 and 29.5 minutes, and 24.5 minutes is the estimate exactly halfway between those bounds.

    minutesmiddle xstaff ff x0 to 10542010 to 20151015020 to 30251435030 to 4035828040 to 50454180total40980give everyone the top of the class1180 divided by 40 = 29.5 minutes
    minutesmiddle xstaff ff x0 to 10542010 to 20151015020 to 30251435030 to 4035828040 to 50454180total40980give everyone the top of the class1180 divided by 40 = 29.5 minutes
    (b) Give everyone the highest time their class allows and the mean is 1180 ÷ 40 = 29.5 minutes, so the true mean lies between 19.5 and 29.5 minutes.

Answer: (a) the modal class is 20 < m ≤ 30 minutes and the estimated mean is 24.5 minutes; (b) the mean is an estimate because each time was replaced by the middle of its class, and the true mean lies between 19.5 and 29.5 minutes

Common mistakes

  • Giving the modal class as 14. That is the frequency of the modal class, not the class itself. The answer to a modal class question is an interval of times, here 20 < m ≤ 30 minutes.
  • Estimating the mean by averaging the five class middles, 5 + 15 + 25 + 35 + 455 = 25. That treats the five classes as equally busy when one of them holds 14 staff and another holds 4. Each middle has to be weighted by its own frequency.

More measuring data problems, worked step by step →

Practice Mean of Grouped Data in the app