← MATH 380 MathScape 0 MATH380

The Distribution of a Variable

14 min read

Jump to a section
Textbook: Moore, McCabe, Craig, Introduction to the Practice of Statistics (IPS), 9th edition  •  Chapter: 1  •  Section: 1

Quick Reference

Field Value
Textbook Moore, McCabe, Craig, Introduction to the Practice of Statistics (IPS), 9th edition
Chapter and section Chapter 1 (Looking at Data: Distributions), Section 1.1 (Data)
Subsection 1.1, “Data: the distribution of a variable”
Course MATH 380 (Statistics)
Pages Page numbers pending faculty verification.

Section 1.1 defines individuals and variables, then names the central object that the rest of Chapter 1 studies: the distribution of a variable. Every display in Section 1.2 (bar graphs, pie charts, histograms, stemplots) and every number in Section 1.3 (center, spread) is a way of describing a distribution, so this definition is the hinge the chapter turns on.


Try This First

Twelve students were asked how many siblings they have. Here are the answers, in the order the students were asked:

2  0  1  3  2  1  0  2  4  1  2  1

Before reading any further, work with this list on paper:

  1. Write down the different values that appear in the list.
  2. Next to each value, write how many students gave that answer.
  3. Which answer was given most often? Which value did no one give above 4?

Hold your small table. You have just built, by hand, the distribution of a variable. That table has a name, and it is the starting point for every graph and summary that follows. (A worked version appears in the first example, so you can check your counts.)

Counting Before Naming

A raw list of data answers a narrow question: what did each individual report. Read straight across, the sibling list is twelve separate facts and not much more. The list becomes informative the moment the question changes from “what did each student say” to “what values show up, and how often.”

That second question is the one worth asking, because it does not depend on the order the students were asked, and it does not grow longer when more students are added. Two values appear: a short collection of possible answers (0, 1, 2, 3, 4) and a count attached to each. A long list of individual responses collapses into a compact summary of values-and-counts.

Think of a variable as a process that takes each individual and gives back a value. The sibling variable takes the first student and returns 2, takes the second and returns 0, and so on. Asking for the distribution runs that process across every individual at once and records not the order of the outputs but the tally of them: how many times each possible output came back. That tally, the values together with how often each one occurs, is the distribution.

This is why the distribution comes before any picture. A bar graph, a histogram, and a stemplot are three different drawings of the same underlying values-and-counts. Build the tally first and each drawing is just a way to show it.

Prerequisite Hub

The distribution is the first object the course studies in depth, so only two short pieces of vocabulary come before it.

Builds on

Unlocks

Once the distribution is the object, every display and summary in Chapter 1 is a way to describe it:

Prerequisite map

The Official Definition

Distribution of a variable. The distribution of a variable tells us what values the variable takes and how often it takes these values.

Source: Moore, McCabe, Craig, Introduction to the Practice of Statistics, 9th edition, Section 1.1 (Data); corroborated by IPS summary sources, where a distribution tells which values occur and how often.

There is no theorem here. The definition has two parts, and naming them separately keeps the idea sharp.

The two parts of every distribution

Part Question it answers Categorical example Quantitative example
The values What can the variable be? The class names (first-year, sophomore, junior, senior) The numbers that occur (0, 1, 2, 3, 4 siblings)
How often How frequently does each value occur? The count or percent in each class The count of individuals at each number

A distribution is incomplete with only one part. A list of possible values with no counts does not say what is common; a pile of counts with no labeled values cannot be read. The distribution is the pairing.

Counts and percents say the same thing two ways

How often a value occurs can be reported as a count (the number of individuals with that value) or as a percent, also called a relative frequency (that count divided by the total, times 100). Counts answer “how many”; percents answer “what share.” Both describe the same distribution, and percents make two distributions of different sizes comparable.

The relative frequency of a value is $$\text{relative frequency} = \frac{\text{count for that value}}{\text{total number of individuals}}.$$

Because every individual lands at exactly one value, the counts add to the total and the relative frequencies add to 1 (the percents add to 100).

Worked Examples

Example 1: Building a quantitative distribution by hand

Take the sibling list from the opener:

2  0  1  3  2  1  0  2  4  1  2  1

Predict first. Scan the list. The answers 1 and 2 seem to show up often; 3 and 4 look rare. Expect a distribution that piles up at 1 and 2 and thins out toward 4.

Now build it. Sweep through the list once and tally each value:

Siblings (value) Tally Count Relative frequency
0 2 students 2 $\frac{2}{12} \approx 0.167$
1 4 students 4 $\frac{4}{12} \approx 0.333$
2 4 students 4 $\frac{4}{12} \approx 0.333$
3 1 student 1 $\frac{1}{12} \approx 0.083$
4 1 student 1 $\frac{1}{12} \approx 0.083$
Total 12 1.000

Check the total. The counts add to $2 + 4 + 4 + 1 + 1 = 12$, which is the number of students, and the relative frequencies add to $1.000$. Both checks pass, so no student was missed or double counted.

Report and compare to the prediction. The most common answer is a tie between 1 and 2 (four students each); 3 and 4 are rare (one student each). The prediction (a pile at 1 and 2, thinning toward 4) matches the built distribution. The table is the distribution of the variable “number of siblings”: the values are 0 through 4, and the counts say how often each occurs.

Example 2: A categorical distribution

A first-year seminar of 20 students records each student’s class standing. The recorded values are:

first-year  first-year  sophomore  first-year  junior
first-year  first-year  sophomore  first-year  senior
first-year  junior      first-year  first-year  first-year
sophomore   first-year  first-year  junior      first-year

Predict first. The variable is class standing, which is categorical (named groups, no numerical order required). “First-year” appears far more than any other label, as expected in a first-year seminar, so its share should be large.

Now build it. Tally each category:

Class standing (value) Count Relative frequency Percent
First-year 13 $\frac{13}{20} = 0.65$ 65%
Sophomore 3 $\frac{3}{20} = 0.15$ 15%
Junior 3 $\frac{3}{20} = 0.15$ 15%
Senior 1 $\frac{1}{20} = 0.05$ 5%
Total 20 1.00 100%

Check the total. Counts: $13 + 3 + 3 + 1 = 20$. Percents: $65 + 15 + 15 + 5 = 100$. Both pass.

Report. The distribution of class standing has four values (first-year, sophomore, junior, senior) with first-year holding 65 percent of the seminar. The prediction (a large first-year share) matches. The same two parts appear as in Example 1: the values are the class names, and “how often” is the count or percent in each.

Example 3: Why percents make two groups comparable

A second seminar of 50 students reports its class standing as 30 first-year, 9 sophomore, 8 junior, and 3 senior. Is the second seminar “more first-year” than the first seminar in Example 2?

Predict first. The second seminar has 30 first-year students against the first seminar’s 13, so counting raw heads, the second seminar has more first-year students. But the seminars are different sizes (50 against 20), so the raw counts alone may mislead. Compare the shares.

Now compute the shares.

Class standing Seminar 1 percent Seminar 2 percent
First-year $\frac{13}{20} = 65\%$ $\frac{30}{50} = 60\%$
Sophomore $\frac{3}{20} = 15\%$ $\frac{9}{50} = 18\%$
Junior $\frac{3}{20} = 15\%$ $\frac{8}{50} = 16\%$
Senior $\frac{1}{20} = 5\%$ $\frac{3}{50} = 6\%$

Report and resolve the prediction. By raw count, Seminar 2 has more than twice as many first-year students (30 against 13). By share, Seminar 1 is slightly more first-year (65 percent against 60 percent). The two readings disagree because the seminars differ in size. The distribution stated as percents answers “what share of this group,” which is what the comparison needs; the distribution stated as counts answers “how many,” which depends on the group size. Both are the same distribution, reported two ways.

Common Misconceptions

Common misconception

the distribution is the list of data values. A common slip is to call the raw list of twelve sibling answers “the distribution.” The list is the data; the distribution is the summary of which values occur and how often. The order of the list does not matter to the distribution, and reordering the twelve answers leaves the same five values with the same five counts. The distribution is the values-and-counts table, not the sequence of individual responses.

Common misconception

the values and the counts are interchangeable. The two parts of a distribution play different roles, and swapping them produces nonsense. In the sibling example, “2” is a value the variable can take (a number of siblings), while “4” is a count (four students reported 1 sibling). Reading the count column as if it were sibling values, or the value column as if it were counts, scrambles the table. Keep “what value” and “how often that value occurs” in separate columns, and label them, so the two are never confused.

Practice Problems

Level 1 Naming the Two Parts

A survey records the eye color of 40 people: brown, blue, green, or hazel. Someone reports “18 of them had brown eyes.”

In the distribution of eye color, which part of the distribution is “brown,” and which part is “18”?

Thought Process

A distribution has two parts: the values the variable takes, and how often each value occurs. Decide which of “brown” and “18” answers “what value” and which answers “how often.”

Show Answer

“Brown” is one of the values the variable (eye color) takes. “18” is how often that value occurs (the count of people with brown eyes). The distribution pairs each value (brown, blue, green, hazel) with its count.

Level 2 Building a Small Distribution

Ten dice rolls came up:

3  6  3  1  4  6  3  2  6  4

Build the distribution of the roll value: list each value that appears and its count. Predict which value is most common before you tally.

Thought Process

Sweep the list once, tallying each value. The values that can appear are 1 through 6, but only the ones that actually occur belong in the table. Check that the counts add to 10.

Show Answer
Roll value Count
1 1
2 1
3 3
4 2
6 3

The value 5 did not occur, so it has count 0 and may be listed with count 0 or omitted. The counts add to $1 + 1 + 3 + 2 + 3 = 10$, matching the ten rolls. The most common values are 3 and 6, tied at three each, so a prediction of “3 or 6” is correct.

Level 3 Counts to Relative Frequencies

A coffee shop sells 200 drinks one morning: 120 coffees, 50 teas, and 30 hot chocolates. Convert this distribution from counts to relative frequencies (as percents), and confirm your percents are consistent.

Thought Process

A relative frequency is the count for a value divided by the total, times 100. The total is 200. After converting all three, the percents should add to 100.

Show Answer
Drink Count Relative frequency
Coffee 120 $\frac{120}{200} = 0.60 = 60\%$
Tea 50 $\frac{50}{200} = 0.25 = 25\%$
Hot chocolate 30 $\frac{30}{200} = 0.15 = 15\%$

Consistency check: $60\% + 25\% + 15\% = 100\%$. The percents add to 100, as they must, because every drink falls into exactly one category.

Level 4 Comparing Two Groups Fairly

Section A has 25 students; 10 of them commute by bus. Section B has 60 students; 18 of them commute by bus. A classmate says “Section B clearly relies on the bus more, because 18 is bigger than 10.”

Use the distribution of “commutes by bus or not” to evaluate the claim.

Thought Process

Raw counts cannot be compared directly when the two groups are different sizes. Convert each section’s bus count to a share of its own section, then compare the shares.

Show Answer

Compute each section’s share of bus commuters:

  • Section A: $\frac{10}{25} = 0.40 = 40\%$.
  • Section B: $\frac{18}{60} = 0.30 = 30\%$.

The claim is not supported. Section B has more bus commuters by raw count (18 against 10), but Section B is more than twice as large. By share, Section A relies on the bus more (40 percent against 30 percent). When the groups differ in size, the comparison must use relative frequencies, not raw counts.

Level 5 Reconstructing a Distribution from Partial Information

A class of 40 students rated a movie on a 1-to-5 scale. The distribution is partly known:

Rating Count
1 2
2 ?
3 10
4 14
5 ?

It is also reported that 25 percent of the class gave a rating of 5. Find the two missing counts, and state the full distribution as relative frequencies.

Thought Process

Two facts pin down the missing counts. First, the count for rating 5 is 25 percent of 40. Second, all five counts must add to 40, the class size. Use the rating-5 percent to find that count, then subtract the known counts from 40 to find the rating-2 count.

Show Answer

Rating 5 count. Twenty-five percent of 40 is $0.25 \times 40 = 10$, so 10 students gave a rating of 5.

Rating 2 count. All counts add to 40: $$2 + (\text{rating 2}) + 10 + 14 + 10 = 40.$$ The known terms sum to $2 + 10 + 14 + 10 = 36$, so the rating-2 count is $40 - 36 = 4$.

Full distribution as relative frequencies:

Rating Count Relative frequency
1 2 $\frac{2}{40} = 0.05 = 5\%$
2 4 $\frac{4}{40} = 0.10 = 10\%$
3 10 $\frac{10}{40} = 0.25 = 25\%$
4 14 $\frac{14}{40} = 0.35 = 35\%$
5 10 $\frac{10}{40} = 0.25 = 25\%$

Check: counts $2 + 4 + 10 + 14 + 10 = 40$, and relative frequencies $5\% + 10\% + 25\% + 35\% + 25\% = 100\%$. Both checks pass, so the reconstructed distribution is consistent.

Resources

Mastery Checklist

Novice (Level 1-2):

Competent (Level 3-4):

Proficient (Level 5):

Connections

Looking back:

Looking ahead:

Real-world connections:



Last updated: 2026-06-16