The Distribution of a Variable
Quick Reference
| Field | Value |
|---|---|
| Textbook | Moore, McCabe, Craig, Introduction to the Practice of Statistics (IPS), 9th edition |
| Chapter and section | Chapter 1 (Looking at Data: Distributions), Section 1.1 (Data) |
| Subsection | 1.1, “Data: the distribution of a variable” |
| Course | MATH 380 (Statistics) |
| Pages | Page numbers pending faculty verification. |
Section 1.1 defines individuals and variables, then names the central object that the rest of Chapter 1 studies: the distribution of a variable. Every display in Section 1.2 (bar graphs, pie charts, histograms, stemplots) and every number in Section 1.3 (center, spread) is a way of describing a distribution, so this definition is the hinge the chapter turns on.
Try This First
Twelve students were asked how many siblings they have. Here are the answers, in the order the students were asked:
2 0 1 3 2 1 0 2 4 1 2 1
Before reading any further, work with this list on paper:
- Write down the different values that appear in the list.
- Next to each value, write how many students gave that answer.
- Which answer was given most often? Which value did no one give above 4?
Hold your small table. You have just built, by hand, the distribution of a variable. That table has a name, and it is the starting point for every graph and summary that follows. (A worked version appears in the first example, so you can check your counts.)
Counting Before Naming
A raw list of data answers a narrow question: what did each individual report. Read straight across, the sibling list is twelve separate facts and not much more. The list becomes informative the moment the question changes from “what did each student say” to “what values show up, and how often.”
That second question is the one worth asking, because it does not depend on the order the students were asked, and it does not grow longer when more students are added. Two values appear: a short collection of possible answers (0, 1, 2, 3, 4) and a count attached to each. A long list of individual responses collapses into a compact summary of values-and-counts.
Think of a variable as a process that takes each individual and gives back a value. The sibling variable takes the first student and returns 2, takes the second and returns 0, and so on. Asking for the distribution runs that process across every individual at once and records not the order of the outputs but the tally of them: how many times each possible output came back. That tally, the values together with how often each one occurs, is the distribution.
This is why the distribution comes before any picture. A bar graph, a histogram, and a stemplot are three different drawings of the same underlying values-and-counts. Build the tally first and each drawing is just a way to show it.
Prerequisite Hub
The distribution is the first object the course studies in depth, so only two short pieces of vocabulary come before it.
Builds on
- Individuals and variables. A distribution is always the distribution of one variable, measured on a set of individuals. The individuals are who is measured; the variable is what is recorded. Without that pair, there is nothing to tally.
- Categorical versus quantitative variables. The kind of variable decides what its distribution looks like and how it is displayed. A categorical variable distributes over named groups; a quantitative variable distributes over numbers.
Unlocks
Once the distribution is the object, every display and summary in Chapter 1 is a way to describe it:
- Displaying categorical data with bar and pie charts
- Histograms
- Stemplots
- Describing a distribution: shape, center, and spread
Prerequisite map
The Official Definition
Distribution of a variable. The distribution of a variable tells us what values the variable takes and how often it takes these values.
Source: Moore, McCabe, Craig, Introduction to the Practice of Statistics, 9th edition, Section 1.1 (Data); corroborated by IPS summary sources, where a distribution tells which values occur and how often.
There is no theorem here. The definition has two parts, and naming them separately keeps the idea sharp.
The two parts of every distribution
| Part | Question it answers | Categorical example | Quantitative example |
|---|---|---|---|
| The values | What can the variable be? | The class names (first-year, sophomore, junior, senior) | The numbers that occur (0, 1, 2, 3, 4 siblings) |
| How often | How frequently does each value occur? | The count or percent in each class | The count of individuals at each number |
A distribution is incomplete with only one part. A list of possible values with no counts does not say what is common; a pile of counts with no labeled values cannot be read. The distribution is the pairing.
Counts and percents say the same thing two ways
How often a value occurs can be reported as a count (the number of individuals with that value) or as a percent, also called a relative frequency (that count divided by the total, times 100). Counts answer “how many”; percents answer “what share.” Both describe the same distribution, and percents make two distributions of different sizes comparable.
The relative frequency of a value is $$\text{relative frequency} = \frac{\text{count for that value}}{\text{total number of individuals}}.$$
Because every individual lands at exactly one value, the counts add to the total and the relative frequencies add to 1 (the percents add to 100).
Worked Examples
Example 1: Building a quantitative distribution by hand
Take the sibling list from the opener:
2 0 1 3 2 1 0 2 4 1 2 1
Predict first. Scan the list. The answers 1 and 2 seem to show up often; 3 and 4 look rare. Expect a distribution that piles up at 1 and 2 and thins out toward 4.
Now build it. Sweep through the list once and tally each value:
| Siblings (value) | Tally | Count | Relative frequency |
|---|---|---|---|
| 0 | 2 students | 2 | $\frac{2}{12} \approx 0.167$ |
| 1 | 4 students | 4 | $\frac{4}{12} \approx 0.333$ |
| 2 | 4 students | 4 | $\frac{4}{12} \approx 0.333$ |
| 3 | 1 student | 1 | $\frac{1}{12} \approx 0.083$ |
| 4 | 1 student | 1 | $\frac{1}{12} \approx 0.083$ |
| Total | 12 | 1.000 |
Check the total. The counts add to $2 + 4 + 4 + 1 + 1 = 12$, which is the number of students, and the relative frequencies add to $1.000$. Both checks pass, so no student was missed or double counted.
Report and compare to the prediction. The most common answer is a tie between 1 and 2 (four students each); 3 and 4 are rare (one student each). The prediction (a pile at 1 and 2, thinning toward 4) matches the built distribution. The table is the distribution of the variable “number of siblings”: the values are 0 through 4, and the counts say how often each occurs.
Example 2: A categorical distribution
A first-year seminar of 20 students records each student’s class standing. The recorded values are:
first-year first-year sophomore first-year junior
first-year first-year sophomore first-year senior
first-year junior first-year first-year first-year
sophomore first-year first-year junior first-year
Predict first. The variable is class standing, which is categorical (named groups, no numerical order required). “First-year” appears far more than any other label, as expected in a first-year seminar, so its share should be large.
Now build it. Tally each category:
| Class standing (value) | Count | Relative frequency | Percent |
|---|---|---|---|
| First-year | 13 | $\frac{13}{20} = 0.65$ | 65% |
| Sophomore | 3 | $\frac{3}{20} = 0.15$ | 15% |
| Junior | 3 | $\frac{3}{20} = 0.15$ | 15% |
| Senior | 1 | $\frac{1}{20} = 0.05$ | 5% |
| Total | 20 | 1.00 | 100% |
Check the total. Counts: $13 + 3 + 3 + 1 = 20$. Percents: $65 + 15 + 15 + 5 = 100$. Both pass.
Report. The distribution of class standing has four values (first-year, sophomore, junior, senior) with first-year holding 65 percent of the seminar. The prediction (a large first-year share) matches. The same two parts appear as in Example 1: the values are the class names, and “how often” is the count or percent in each.
Example 3: Why percents make two groups comparable
A second seminar of 50 students reports its class standing as 30 first-year, 9 sophomore, 8 junior, and 3 senior. Is the second seminar “more first-year” than the first seminar in Example 2?
Predict first. The second seminar has 30 first-year students against the first seminar’s 13, so counting raw heads, the second seminar has more first-year students. But the seminars are different sizes (50 against 20), so the raw counts alone may mislead. Compare the shares.
Now compute the shares.
| Class standing | Seminar 1 percent | Seminar 2 percent |
|---|---|---|
| First-year | $\frac{13}{20} = 65\%$ | $\frac{30}{50} = 60\%$ |
| Sophomore | $\frac{3}{20} = 15\%$ | $\frac{9}{50} = 18\%$ |
| Junior | $\frac{3}{20} = 15\%$ | $\frac{8}{50} = 16\%$ |
| Senior | $\frac{1}{20} = 5\%$ | $\frac{3}{50} = 6\%$ |
Report and resolve the prediction. By raw count, Seminar 2 has more than twice as many first-year students (30 against 13). By share, Seminar 1 is slightly more first-year (65 percent against 60 percent). The two readings disagree because the seminars differ in size. The distribution stated as percents answers “what share of this group,” which is what the comparison needs; the distribution stated as counts answers “how many,” which depends on the group size. Both are the same distribution, reported two ways.
Common Misconceptions
the distribution is the list of data values. A common slip is to call the raw list of twelve sibling answers “the distribution.” The list is the data; the distribution is the summary of which values occur and how often. The order of the list does not matter to the distribution, and reordering the twelve answers leaves the same five values with the same five counts. The distribution is the values-and-counts table, not the sequence of individual responses.
the values and the counts are interchangeable. The two parts of a distribution play different roles, and swapping them produces nonsense. In the sibling example, “2” is a value the variable can take (a number of siblings), while “4” is a count (four students reported 1 sibling). Reading the count column as if it were sibling values, or the value column as if it were counts, scrambles the table. Keep “what value” and “how often that value occurs” in separate columns, and label them, so the two are never confused.
Practice Problems
A survey records the eye color of 40 people: brown, blue, green, or hazel. Someone reports “18 of them had brown eyes.”
In the distribution of eye color, which part of the distribution is “brown,” and which part is “18”?
Ten dice rolls came up:
3 6 3 1 4 6 3 2 6 4
Build the distribution of the roll value: list each value that appears and its count. Predict which value is most common before you tally.
A coffee shop sells 200 drinks one morning: 120 coffees, 50 teas, and 30 hot chocolates. Convert this distribution from counts to relative frequencies (as percents), and confirm your percents are consistent.
Section A has 25 students; 10 of them commute by bus. Section B has 60 students; 18 of them commute by bus. A classmate says “Section B clearly relies on the bus more, because 18 is bigger than 10.”
Use the distribution of “commutes by bus or not” to evaluate the claim.
A class of 40 students rated a movie on a 1-to-5 scale. The distribution is partly known:
| Rating | Count |
|---|---|
| 1 | 2 |
| 2 | ? |
| 3 | 10 |
| 4 | 14 |
| 5 | ? |
It is also reported that 25 percent of the class gave a rating of 5. Find the two missing counts, and state the full distribution as relative frequencies.
Resources
- Textbook section. Moore, McCabe, Craig, Introduction to the Practice of Statistics (IPS), 9th edition, Chapter 1, Section 1.1 (Data), subsection “Data: the distribution of a variable.” Page numbers pending faculty verification.
- OpenStax Introductory Statistics 2e, Chapter 2 Key Terms. The open-text definitions of frequency and relative frequency that match the “how often” part of a distribution: openstax.org/books/introductory-statistics-2e/pages/2-key-terms.
Mastery Checklist
Novice (Level 1-2):
Competent (Level 3-4):
Proficient (Level 5):
Connections
Looking back:
- Individuals and variables supplies the who (individuals) and the what (the variable) that a distribution summarizes.
- Categorical versus quantitative variables decides whether a distribution spreads over named groups or over numbers, which sets the right display.
Looking ahead:
- Bar and pie charts draw the distribution of a categorical variable.
- Histograms and stemplots draw the distribution of a quantitative variable.
- Shape, center, and spread describes a distribution in words and numbers, once it has been displayed.
Real-world connections:
- A poll result (“42 percent favor the proposal”) is one value-and-share from the distribution of a survey variable.
- A weather summary of “days of rain per month” is the distribution of a count variable across the year.
- A grade report of how many students earned each letter grade is the distribution of the grade variable.
| Previous | Up | Next |
|---|---|---|
| Categorical vs Quantitative Variables | Skills Index | Bar and Pie Charts |
Last updated: 2026-06-16