Describing a Distribution: Shape, Center, and Spread
Quick Reference
Textbook: Moore, McCabe, and Craig, Introduction to the Practice of Statistics (IPS), 9th edition. Location: Chapter 1 (Looking at Data: Distributions), Section 1.2 (Displaying Distributions with Graphs), subsection “Interpreting histograms: overall pattern (shape, center, spread) and outliers.” Pages: Page numbers pending faculty verification. Role in the course: The reading skill that turns a finished graph into words. Every later summary in the course (mean and median, quartiles and the interquartile range, density curves and the Normal model) answers a question that this skill first asks in plain language. A student who can name shape, center, and spread can choose the right numerical summary in the next section.
Try This First
📋 A two-minute warmup before any definition (click to open)
The histogram below shows the number of text messages a sample of 40 students sent in one day. Each block is one bar of the histogram, drawn here with characters so it reads on any screen.
Count of students
10 | ####
8 | #######
6 | #######
4 | ##########
2 | ##############
0 +----------------------------
0 20 40 60 80 100 120
Messages sent in one day
Answer three questions in plain words before reading any further.
- Where do most students fall, roughly, on the horizontal axis?
- Does the picture look balanced left to right, or does one side stretch out?
- Is there any student who sits far away from the rest?
Compare your answers
- Most students cluster between about 0 and 40 messages. That cluster is the center of the picture.
- The right side stretches out toward 120 while the left side stops at 0. The picture is not balanced; it leans with a long right tail.
- The handful of students out near 100 to 120 sit far from the main cluster. They are candidates for outliers.
You just described shape, center, spread, and a possible outlier without a single formula. The section below names each of those ideas precisely.
A histogram or a stemplot is finished, and the next task is to say what it shows. The eye already does most of the work: it finds the hump where the data pile up, it notices whether one tail runs longer than the other, and it flags any point sitting off by itself. Putting those observations into a fixed vocabulary is the entire skill.
The fixed vocabulary has three words plus one warning. The three words are shape, center, and spread. The warning is to watch for an outlier, a value that falls outside the overall pattern. Reading a distribution always proceeds in that order: find the overall pattern first (shape, center, spread), then look for striking deviations from it (outliers). Naming the pattern before hunting for exceptions keeps a single unusual point from distorting the description of the whole.
Prerequisite Hub
Builds on
| Skill | Why it comes first |
|---|---|
| Histograms | The overall pattern is read off a histogram. The graph must exist before its shape can be described. |
| Stemplots | A stemplot is the second display whose shape, center, and spread are described, useful for a small data set. |
| Distribution of a Variable | A distribution records which values a variable takes and how often. Its pattern is exactly what this skill puts into words. |
Unlocks
| Skill | What this enables |
|---|---|
| Measures of Center: Mean and Median | Turns the eyeballed center into a number, with the choice between mean and median driven by the shape named here. |
| Measures of Spread: Quartiles and IQR | Turns the eyeballed spread into numbers and gives a rule for flagging outliers. |
| Density Curves | Replaces the histogram with a smooth curve whose shape, center, and spread carry the same meaning. |
Cross-course prerequisites: none.
Prerequisite Map
Quick Reference
| Property | Value |
|---|---|
| Chapter.Section | IPS 1.2 |
| Course | MATH380 |
| Difficulty | Introductory |
| Time | ~25 minutes |
The Three Words at a Glance
| Word | Question it answers | What to look at | Plain-language report |
|---|---|---|---|
| Shape | Is the picture balanced or lopsided? How many humps? | The outline of the bars | Symmetric, skewed right, skewed left, single peak, two peaks |
| Center | Where is a typical value? | The horizontal location of the middle of the data | “Most values are near 35” |
| Spread | How far apart are the values? | The horizontal width from smallest to largest | “Values run from 0 to 120” |
| Outlier (the warning) | Does any value sit outside the pattern? | Points far from the main cluster | “One value near 250 stands apart” |
Reading the Lean of a Distribution
| Shape | What the tails do | Memory aid |
|---|---|---|
| Symmetric | Left and right sides are approximate mirror images | A balanced seesaw |
| Skewed to the right | The right tail (larger values) stretches farther | The tail points the way it is named: toward the right |
| Skewed to the left | The left tail (smaller values) stretches farther | The tail points the way it is named: toward the left |
Official Definitions
Overall pattern of a distribution (shape, center, spread).
In any graph of data, look for the overall pattern and for striking deviations from that pattern. You can describe the overall pattern of a distribution by its shape, center, and spread. An important kind of deviation is an outlier, an individual value that falls outside the overall pattern.
Source: Moore, McCabe, Craig, IPS 9th edition, Section 1.2 (Displaying Distributions with Graphs: examining a distribution).
Symmetric and skewed distributions.
A distribution is symmetric if the right and left sides of its graph are approximately mirror images of each other. A distribution is skewed to the right if the right side of the graph (containing the half of the observations with larger values) extends much farther out than the left side; it is skewed to the left if the left side extends much farther out than the right side.
Source: Moore, McCabe, Craig, IPS 9th edition, Section 1.2 (Displaying Distributions with Graphs: examining a distribution).
Read the two definitions together. The first definition fixes the order of operations for the eye: describe the overall pattern, then mark the deviations. It also names the three reporting words and defines an outlier as a value outside the pattern, not merely the largest or smallest value. The second definition turns “shape” from a vague impression into a decision about the tails: the side that extends much farther names the skew. The word “approximately” appears on purpose. Real data rarely produce a perfect mirror image, so a distribution is called symmetric when the two sides are close, not identical.
There is no theorem in this section. The two definitions are the full toolkit, and the skill is applying the tail comparison and the outside-the-pattern test correctly to real graphs.
Worked Examples
Example 1: Naming shape, center, and spread from a stemplot
A clinic records the resting heart rate, in beats per minute, of 20 patients. The stemplot below uses tens as the stem and ones as the leaf, so the row 6 | 2 5 8 represents 62, 65, and 68.
5 | 8 9
6 | 2 4 5 7 8 9
7 | 0 1 2 3 5 6 7 8
8 | 1 2 4
9 | 0
Predict first. Before reading on, predict in one phrase where the center sits and whether the picture is balanced. Write your phrase down.
Read the shape. The longest rows are the 60s and 70s. Above the 70s the rows shrink (three values in the 80s, one in the 90s), and below the 60s there are only two values in the 50s. Both tails are short and the bulk sits in the middle two stems, so the picture is roughly symmetric with a single peak, leaning very slightly toward the higher values.
Read the center. The middle of the 20 ordered values lands in the low 70s. A reasonable plain-language center is “a typical resting heart rate is about 70 to 72 beats per minute.” (The next section replaces this estimate with the median.)
Read the spread. The smallest value is 58 and the largest is 90, so the data run across a width of
$$90 - 58 = 32 \text{ beats per minute}.$$
Check against your prediction. A complete description reads: the distribution of resting heart rates is roughly symmetric with a single peak, centered near 70 to 72 beats per minute, and spread from 58 to 90 beats per minute, with no value sitting far from the rest. If the prediction matched the location of the center and the balance of the tails, the reading is confirmed.
Example 2: Skewed to the right, with an outlier check
A real estate report lists the sale prices, in thousands of dollars, of 15 homes in a neighborhood:
$$185,\ 190,\ 195,\ 200,\ 205,\ 210,\ 215,\ 220,\ 225,\ 230,\ 235,\ 240,\ 250,\ 260,\ 640.$$
Predict first. Glance at the list. Predict which side will have the longer tail, and predict whether any single value stands apart.
Read the shape. Fourteen of the fifteen prices sit between 185 and 260, a tight cluster. The fifteenth price, 640, sits far above that cluster. The right side stretches much farther than the left, so the distribution is skewed to the right.
Read the center. Ignoring the lone high price for the moment, a typical home sells for roughly 215 to 220 thousand dollars. That cluster is the overall pattern.
Read the spread of the main pattern. Within the cluster, prices run from 185 to 260, a width of
$$260 - 185 = 75 \text{ thousand dollars}.$$
Mark the deviation. The value 640 falls well outside the overall pattern of the other fourteen homes. It is a candidate outlier, worth investigating (a mansion, a data-entry error, or a luxury sale). The point of describing the pattern first is precisely this: the cluster defines what “outside the pattern” means.
Conclusion. The distribution of sale prices is skewed to the right, centered near 215 to 220 thousand dollars, with the main group spread from 185 to 260 thousand dollars, and one value near 640 thousand dollars that falls outside the pattern.
Example 3: Two peaks reveal two groups
A fitness center records the times, in minutes, that 24 members spent on a treadmill in one session. A histogram of the times shows the following bar counts.
| Time interval (min) | Count |
|---|---|
| 0 to 10 | 2 |
| 10 to 20 | 7 |
| 20 to 30 | 4 |
| 30 to 40 | 1 |
| 40 to 50 | 3 |
| 50 to 60 | 6 |
| 60 to 70 | 1 |
Predict first. Predict how many humps the histogram has, and predict what a single “typical” time would miss.
Read the shape. The counts rise to a peak at 10 to 20 minutes, fall to a low of 1 in the 30 to 40 interval, then rise again to a second peak at 50 to 60 minutes. Two separated humps appear, so the distribution has two peaks (it is bimodal). The dip between them at 30 to 40 minutes is the key feature.
Read the center. Reporting one center here would be misleading. A single “typical time” near 35 minutes lands in the dip, where almost no member actually exercised. Two peaks usually signal two different groups mixed together (perhaps a short-warmup group and a long-workout group), so the honest description names both clusters rather than averaging across them.
Read the spread. Times run from 0 to 70 minutes, a width of
$$70 - 0 = 70 \text{ minutes}.$$
Conclusion. The distribution of treadmill times has two peaks, one near 10 to 20 minutes and one near 50 to 60 minutes, separated by a dip near 30 to 40 minutes, with values spread from 0 to 70 minutes. The two peaks are the headline: a single center would hide the structure that matters.
Common Misconceptions
the direction of the skew is the side where the tall bars are. The skew is named for the side with the long tail, not the side with the peak. In a right-skewed distribution the tall bars sit on the left and a thin tail runs to the right, yet the distribution is called skewed to the right. The tail points the way the distribution is named. Predict-then-check: in Example 2 the tall cluster is on the left near 200, and the lone high price stretches to the right, so the distribution is skewed to the right even though the peak is on the left. Reading the peak instead of the tail reverses the answer.
a histogram is a picture of the objects, so a tall bar means a large value. A bar is tall when many individuals fall in that interval, not when the values there are large. The height of a bar counts individuals; the horizontal position of the bar reports the value. In Example 1, the tall rows are the 60s and 70s because most patients have heart rates there, not because those heart rates are large. Confusing the height of a bar with the size of the value turns the count axis into a value axis and breaks every reading. Check on a small case: a histogram of exam scores with a tall bar at “50 to 60” means many students scored in the 50s, not that the scores there are high.
every distribution has exactly one center worth reporting. When a distribution has two peaks, a single center lands in the dip between the groups and describes almost no one. In Example 3 a single “typical” time near 35 minutes falls in the valley where the fewest members exercised. Two peaks usually mean two groups are mixed, so the correct description names both clusters and the gap between them. Forcing one center onto a two-peaked picture hides the most important feature.
the largest value is automatically an outlier. An outlier is a value outside the overall pattern, not simply the maximum. In a symmetric distribution the largest value can sit snugly at the end of a smooth, gradual tail and belong to the pattern. A value counts as an outlier only when a clear gap separates it from the rest, as the 640 does in Example 2. Describe the overall pattern first, then judge whether a point falls outside it; calling the maximum an outlier by reflex skips that judgment.
Practice Problems
For each described distribution, state whether it is symmetric, skewed to the right, or skewed to the left.
(a) A histogram of household incomes where most households cluster at lower incomes and a thin tail of very high incomes runs to the right.
(b) A histogram of adult heights that is balanced left to right around a single central peak.
(c) A histogram of exam scores on an easy test where most students score high and a thin tail of low scores runs to the left.
A stemplot of the ages of 12 people at a community event is shown, with tens as the stem and ones as the leaf, so 2 | 1 4 represents 21 and 24.
1 | 8 9
2 | 1 4 6 7
3 | 0 3 5
4 | 2 8
6 | 5
State the center in plain words and give the spread as the difference between the largest and smallest values.
A histogram of the number of days customers waited for a refund has tall bars near 2 to 5 days and a thin tail of long waits stretching out to 40 days. A classmate says the distribution is skewed to the left because the tall bars are on the left.
Which statement correctly describes the distribution?
(A) Skewed to the left, because the tall bars are on the left.
(B) Skewed to the right, because the long tail runs to the right.
(C) Symmetric, because there is one clear peak.
(D) There is not enough information to decide the shape.
A histogram of quiz scores has its tallest bar over the interval 60 to 70. A student writes: “The tallest bar is the largest score, so the highest quiz score is about 65.” Evaluate the claim and state what the tallest bar actually reports.
A histogram of the number of items in 30 shopping carts shows bar counts as follows.
| Items in cart | Count of carts |
|---|---|
| 0 to 5 | 9 |
| 5 to 10 | 8 |
| 10 to 15 | 5 |
| 15 to 20 | 3 |
| 20 to 25 | 2 |
| 25 to 30 | 1 |
| 45 to 50 | 1 |
| 50 to 55 | 1 |
(a) Describe the distribution in one sentence using shape, center, and spread.
(b) Identify any value that falls outside the overall pattern and justify the choice by referring to the pattern, not just the size of the value.
Mental Model
The four-word reading checklist.
For any histogram or stemplot, walk through four prompts in order.
- Shape. Is the picture balanced (symmetric) or does one tail run longer (skewed, named for the long tail)? How many humps?
- Center. Where on the horizontal axis does a typical value sit?
- Spread. How wide is the horizontal range from smallest to largest?
- Outliers. After naming the pattern, does any value sit outside it across a clear gap?
Pattern first, exceptions last. Describing shape, center, and spread before hunting for outliers keeps one unusual point from rewriting the description of the whole distribution.
Mastery Checklist
Novice (Level 1-2):
Competent (Level 3-4):
Proficient (Level 5):
Connections
Looking back:
- Histograms and Stemplots build the displays whose pattern this skill puts into words.
- Distribution of a Variable defines the pattern of values that shape, center, and spread summarize.
Looking ahead:
- Measures of Center: Mean and Median replaces the eyeballed center with a number, and the shape named here decides which number to report.
- Measures of Spread: Quartiles and IQR replaces the eyeballed spread with numbers and gives a rule for flagging outliers.
- Density Curves smooths the histogram into a curve that carries the same shape, center, and spread.
Real-world connections:
- A salary report that ignores the right skew of incomes and reports only an average overstates what a typical worker earns.
- A quality-control histogram with two peaks signals two production lines mixed together, a structure a single average would hide.
Resources
| Resource | Where it points |
|---|---|
| IPS 9th edition, Section 1.2 (Displaying Distributions with Graphs) | Moore, McCabe, Craig, IPS 9e, Chapter 1, Section 1.2. The primary source for the overall-pattern definition and the symmetric-versus-skewed definitions in this skill. |
| OpenStax Introductory Statistics 2e, Chapter 2 introduction (shape of a distribution) | openstax.org. Free open companion treatment of shape, center, and spread. |
| Previous | Up | Next |
|---|---|---|
| Distribution of a Variable | Skills Index | Measures of Center: Mean and Median |
Last updated: 2026-06-16