Comparing Data Sets

Averages and spreads together, never alone.

The same average

Two football teams record the points they score in each game of a season. Team A’s scores have a median of 20 points, a lower quartile of 18 and an upper quartile of 22. Team B’s scores also have a median of 20 points, but a lower quartile of 10 and an upper quartile of 30.

By the median alone, the two teams are the same. The difference shows in the spread. Team A’s interquartile range is 22 − 18 = 4 points, so the middle half of its games all ended within 4 points of each other. Team B’s interquartile range is 30 − 10 = 20 points, five times as wide.

team A010203040

Team A’s box runs from 18 to 22, around a median of 20: its scores are very consistent.

team Ateam B010203040

On the same scale, team B’s median is also 20, but its box stretches from 10 to 30.

One average and one spread, in a sentence

A comparison of two data sets needs two parts: an average, which says where each set is centered, and a measure of spread, which says how consistent its values are. Give both numbers for both sets, and then say what they mean in the context of the data.

For the two teams: “The teams have the same median score, 20 points. Team A’s interquartile range is 4 points and team B’s is 20 points, so team A’s scores are much more consistent.”

A list of numbers is not yet a comparison. “Team A’s IQR is 4 and team B’s is 20” is true, but it does not say which team is more consistent, or what the numbers are measuring. Each number needs a sentence that says what it shows about the teams.

When the averages differ too

Usually both parts differ. Two classes take the same test. Class A has a median of 62 marks and an interquartile range of 10 marks. Class B has a median of 55 marks and an interquartile range of 25 marks.

The comparison makes one statement about each measure. “On average, Class A scored higher: its median is 62 marks, 7 more than Class B’s 55. Class A’s scores were also more consistent: the middle half of its marks spans 10 marks, while Class B’s spans 25.”

The average and the spread answer different questions, so neither one decides the other. A class can have the higher average and the wider spread. And a larger spread does not mean a class did worse; it means its marks were less alike.

class Aclass B2030405060708090

Class A’s median, 62, is farther to the right than Class B’s, 55, so Class A scored higher on average. Class A’s box, 10 marks wide, is narrower than Class B’s, 25 marks wide, so its marks were more consistent.

Higher is not always better

Say which way is better in the context. In a test, a higher median means the class did better. In a race, a higher median time means the runners were slower. For two runners with median times of 12.4 seconds and 12.9 seconds over 100 meters, the comparison reads: “On average the first runner was faster, by 0.5 seconds.”

Which average and which spread?

The two measures come in pairs. The median goes with the interquartile range: both are found from the positions of the values once they are in order. The mean goes with the standard deviation: both use the size of every value. Quote the same pair for both data sets, so that like is compared with like.

When a data set has an outlier, use the median and the interquartile range. An outlier moves the mean, the range and the standard deviation a long way, but barely moves the median and the interquartile range.

An outlier decides the choice

Pat and Sam record how many minutes their journey to work takes on seven days. Pat’s times, in order, are 12, 14, 15, 15, 16, 17 and 60 minutes: on the last day a train broke down. Sam’s times are 14, 16, 17, 18, 19, 20 and 22 minutes.

Pat’s mean is 149 ÷ 7 ≈ 21.3 minutes and Sam’s is 126 ÷ 7 = 18 minutes, so the means say Pat’s journey takes longer. But on six of the seven days Pat took 17 minutes or less, which is less than Sam’s median. The one 60-minute day lifts Pat’s mean above every other time Pat recorded.

The medians describe an ordinary day. Pat’s median is 15 minutes and Sam’s is 18 minutes. For the spread, Pat’s lower quartile is 14 and upper quartile 17, an interquartile range of 3 minutes; Sam’s quartiles are 16 and 20, an interquartile range of 4 minutes. The range and the standard deviation tell another story: Pat’s range is 60 − 12 = 48 minutes and Sam’s is 22 − 14 = 8, and Pat’s standard deviation, about 15.9 minutes, is more than six times Sam’s, about 2.4 minutes. Both are measuring the one bad day.

The 60-minute day is an outlier: the upper limit is 17 + 1.5 × 3 = 21.5 minutes, and 60 is far above it. So the comparison uses the median and the IQR, and mentions the outlier on its own: “On a typical day Pat’s journey is shorter, with a median of 15 minutes against Sam’s 18, and Pat’s times are slightly more consistent, with an IQR of 3 minutes against 4. Pat had one very long journey, of 60 minutes, when a train broke down.”

PatSam102030405060

Pat’s box, from 14 to 17, is narrower than Sam’s, from 16 to 20, and sits farther to the left. Only the long whisker out to 60, one day, makes Pat’s times look spread out.

When the mean and the standard deviation are the right pair

When neither set has an outlier and the values are roughly symmetric, the mean and the standard deviation are the better pair, because they use every value. Two machines filling bottles are usually compared this way: “Both machines fill bottles with a mean of 500 milliliters, but machine A’s standard deviation is 2 milliliters and machine B’s is 6 milliliters, so machine A is more consistent.”

Practice Comparing Data Sets in the app