Two kinds of answer
Ask everyone in a class a question and write down every answer: the answers are data. Some questions are answered with a name. What is your favorite pet? Cat, dog, fish. Each answer puts a person into a group, called a category, so these answers are categorical data.
Other questions are answered with an amount on a number scale. How tall are you? 152 cm. How many goals did you score? 3. These answers are numerical data. Some are measurements, such as a height, and some are counts, such as a number of goals; both can be put in order, added and averaged.
Categories are counted
Fifteen children name their favorite pet: 7 choose a cat, 5 a dog and 3 a fish. Each category can be counted. The count for each category is its frequency, and the frequencies add up to everyone asked: 7 + 5 + 3 = 15. The most common answer is cat.
A category cannot be averaged. A mean adds up the answers and shares the total out equally, and cat + dog + fish has no total: a name has no size. The mean of the fifteen answers does not exist.
Each bar counts the children in one category: 7 chose a cat, 5 a dog and 3 a fish.
Numbers can be averaged
Eight students measure their heights, in centimeters: 146, 148, 150, 150, 151, 152, 153 and 154. These are numbers on a scale, so they can be put in order, subtracted and averaged.
The total is 1204 cm, and the mean is 1204 ÷ 8 = 150.5 cm. That mean is a height, and it means something: it lies between the shortest student, 146 cm, and the tallest, 154 cm, and it is the height each of them would have if the total were shared out equally.
One dot for each student, above that student’s height in centimeters. The mean, 150.5 cm, lies between the shortest height and the tallest.
Counts and measurements
Numerical data come in two kinds. A count, such as the number of children in a family, can only be a whole number: a family has 2 children or 3, never 2.4. Data like these are discrete: they take separate values, with gaps between them.
A measurement, such as a height, can take any value in a range. Between 150 cm and 151 cm there are 150.2 cm, 150.25 cm and every other length. Data like these are continuous. A height written as 150 cm has been rounded to the nearest centimeter, so it stands for every height from 149.5 cm up to 150.5 cm.
The mean of discrete data need not be a whole number. Five families with 2, 3, 2, 3 and 2 children have 12 children in all, and 12 ÷ 5 = 2.4 children for each family. No family has 2.4 children, but the mean still describes the five families together.
Numbers that are names
Some numbers are labels. A bus route numbered 12, a player wearing the number 7 and a room numbered 105 are named by their numbers, not measured by them. Route 12 is not twice as much as route 6, and the mean of the numbers on a team’s shirts tells you nothing. Data like these are categorical, even though they are written in digits.
The test is whether arithmetic makes sense. If adding two answers or finding their mean means nothing, the data are categorical.
The type of data chooses the chart
Categorical data are drawn as a bar chart, with one bar for each category and gaps between the bars because nothing lies between one category and the next, or as a pie chart, which shows each category’s share of the whole.
Discrete data are drawn on a number scale, as a dot diagram or a vertical line chart. Each whole number has its own column of dots or its own line, and nothing is drawn between the whole numbers, because no value lies there. A small set of measurements rounded to whole units, like the eight heights above, can be drawn as a dot diagram too.
Larger sets of continuous data are first grouped into classes, which are ranges of values, and then drawn as a histogram. A chart that does not match its data misleads: a pie chart of daily temperatures, for example, treats the temperatures as shares of a whole, and they are not.
Grouping measurements into classes
Measurements rarely repeat exactly, so they are grouped before they are drawn. The class holds every value x from 12 up to, but not including, 14. So 12.0 is in that class and 14.0 is not. The next class, , starts exactly where the last one stops, so every value belongs to exactly one class.
A histogram draws one bar over each class, as tall as the number of values in it. The bars touch, because the classes meet with no gap between them.
Worked example: Three Sets of Fitness Data and the Display That Suits Each
Question A PE teacher has three sets of data about one class: each student's favorite sport, the number of pull-ups each student can do, and each student's time for a 2.4 km run. (a) Choose a suitable display for each set of data, and give the reason from the type of data. (b) The run times, in minutes, of 12 students are 11.5, 12.0, 13.2, 13.9, 14.0, 14.6, 12.8, 15.1, 16.4, 13.5, 14.0 and 17.2. They are grouped into the classes 10 ≤ x < 12, 12 ≤ x < 14, 14 ≤ x < 16 and 16 ≤ x < 18 for a histogram. How many students are in the class 12 ≤ x < 14, and in which class does a time of 14.0 minutes belong?
1.A favorite sport is a name, so the data are categorical. A bar chart suits them: there is one bar for each sport, and the gaps between the bars show that the categories are separate. A pie chart would also do.
A favorite sport is a name, so the data are categorical, and a bar chart has one separate bar for each sport. 2.A number of pull-ups is a count, 0, 1, 2 and so on, so the data are discrete. A dot diagram or a vertical line chart suits them, because each whole number has its own column of dots or its own line, and nothing is drawn between the whole numbers.
A number of pull-ups is a count, so the data are discrete, and a dot diagram has a column of dots at each whole number. 3.(a) A run time is a measurement and can be any value, such as 13.27 minutes, so the data are continuous. A histogram suits them: the times are grouped into classes, and the bars touch because each class starts where the class before it ends.
(a) A run time is a measurement, so the data are continuous, and a histogram groups the times into classes whose bars touch. 4.Tally each time into its class. The class 12 ≤ x < 14 includes 12.0 but does not include 14.0, so a time of 14.0 goes into 14 ≤ x < 16. The tallies are 1, 5, 4 and 2. Check: 1 + 5 + 4 + 2 = 12 students.
Tally each time into its class. 14.0 is left out of 12 ≤ x < 14 and goes into 14 ≤ x < 16. The tallies are 1, 5, 4 and 2. 5.(b) 5 students are in the class 12 ≤ x < 14: their times are 12.0, 13.2, 13.9, 12.8 and 13.5. Each time of 14.0 minutes belongs to the class 14 ≤ x < 16.
(b) 5 students are in the class 12 ≤ x < 14, and 14.0 belongs to 14 ≤ x < 16.
Answer: (a) a bar chart for the sports, which are categorical; a dot diagram or a vertical line chart for the pull-ups, which are discrete; a histogram for the run times, which are continuous; (b) 5 students, and 14.0 belongs to 14 ≤ x < 16
Common mistakes
- Drawing a histogram for the favorite sports. A histogram needs a number scale along its horizontal axis, and the sports have no order and no scale. Separate bars in a bar chart are correct for categories.
- Counting a time of 14.0 minutes in both classes, or in 12 ≤ x < 14. The sign < at the upper end means that 14 is left out of that class, and ≤ at the lower end of the next class means that it is included there. Every value belongs to exactly one class.