Prefer to click through the interactive model?

Distributions — study guide

The same fragments the interactive model serves, read in order. One source, two views.

The shape of a pile of numbers

Measure the same thing many times and you get a pile of numbers: the height of everyone in a room, the time each request took, the distance a sensor reported on every reading. A distribution is just what that pile looks like when you set all the numbers down together and step back — not any single value, but the shape the whole batch makes.

Three questions describe that shape, and almost everything else is detail. Where is the center — what counts as a typical value? How spread out are the numbers — do they huddle close to the center or scatter wide? And are there any strays — a few values sitting far from the rest that deserve a second look. Answer those three and you can describe a batch you have never plotted to someone who has never seen it.

The reason to describe the shape rather than list the numbers is that the shape is what you can actually reason about. A thousand raw request times tell you nothing at a glance; the sentence that they center near a fifth of a second, cluster tightly, and hide two very slow outliers tells you almost everything. The rest of this concept is those three questions, one at a time.

Three ways to say typical

mode median mean (balance point) all three meet at the center
On a symmetric pile of numbers, the mean, the median, and the mode all fall at the center.

There is no single answer to where's the center, because typical means at least three different things. The mean is the arithmetic average: add every value and divide by how many there are. It is the balance point of the pile — the spot where the numbers would sit level if you laid them on a see-saw like the balance point of a see-saw: the mean is the spot where the numbers would sit level, while the median is just whoever is standing in the middle of the line. Because the mean uses every value, every value gets a vote in where it lands.

The median is the middle value: line the numbers up in order and take the one in the middle (or the average of the two middle ones when the count is even). It ignores how far away the extremes are and cares only about their position, so half the batch sits below it and half above. The mode is the value that shows up most often — the tallest stack when equal values pile on top of each other. A batch can have one mode, several, or none worth naming.

These three agree only when the pile is roughly symmetric. Pull one side out of shape and they separate, and which one you should quote depends on why you are asking: the mean for a fair share of a total, the median for a typical member, the mode for the most common case. Reaching for the mean out of habit is the most common way a summary quietly misleads.

How tightly the numbers cluster

same center tight wide Q1 Q3 IQR: middle half
Two batches share one center but differ in spread; the interquartile range is the width of the middle half between Q1 and Q3.

Two batches can share the same center and still be nothing alike. One might be a tight huddle a hair's width from the middle; the other a loose crowd sprawled across a wide range like the difference between a tight huddle and a scattered crowd: same middle, but one gives the numbers far more elbow room than the other. Spread is the measure of that elbow room — how far the numbers sit from the center — and a center quoted without it hides half the story.

The crudest measure is the range: the largest value minus the smallest. It is easy but fragile, since it depends entirely on the two most extreme numbers. Variance and standard deviation are the workhorses instead: variance is the average of the squared distances from the mean, and the standard deviation is its square root, which puts the measure back in the original units. A larger standard deviation means a wider, looser pile. One wrinkle worth knowing: when your numbers are a sample meant to stand in for a larger group, the standard formula divides by one less than the count (n minus 1) rather than the count itself, which nudges the estimate upward to correct a known bias.

The interquartile range takes a sturdier path. Sort the numbers and find the quartiles: the first quartile (Q1) is the value a quarter of the way in, the third quartile (Q3) three quarters of the way in. The interquartile range is simply Q3 minus Q1 — the width of the middle half of the data. Because it throws away the extreme quarter on each end, it describes the bulk of the pile without letting a single wild value stretch it.

The strays that pull the mean

outlier median (holds) mean pulled Q3 + 1.5 x IQR fence beyond the fence, so flagged
A single far-out value drags the mean toward it while the median holds, and the value falls beyond the upper 1.5 times IQR fence.

An outlier is a value sitting far from the rest of the pile — one request that took forty times longer than the others, one sensor reading that spiked when nothing moved. Outliers matter out of proportion to their count, because the mean feels every one of them. Since the average uses every value, a single extreme number drags it toward itself like one very tall visitor walking into a room: the average height jumps, but the person standing in the middle of the line is unchanged, while the median, which cares only about position, barely twitches. Watching the mean and the median pull apart is often the first sign a stray is present.

That sensitivity is exactly why you compare the two. If the mean sits well above the median, a few large values are inflating the average; if it sits well below, a few small ones are dragging it down. The median and the interquartile range are called robust for this reason — they describe the bulk of the data without letting a lone extreme rewrite the summary.

A common rule of thumb draws the line automatically. Compute the interquartile range, then treat anything below Q1 minus one-and-a-half times that range, or above Q3 plus one-and-a-half times it, as a candidate outlier. The rule only flags values for a look; it does not decide their fate. An outlier can be a typo, a broken sensor, or the single most important event in the batch, and only knowing where the numbers came from tells you which.

When the pile leans

left skew mean below median symmetric mean = median right skew mean above median accent line = mean, plain line = median
In a skewed pile the mean is pulled toward the long tail relative to the median; in a symmetric pile the two coincide.

A pile of numbers need not be symmetric. When it is, the two sides mirror each other and the mean and median land together in the middle. Skew is the name for the lean when they do not — when one tail stretches out far longer than the other, dragging part of the pile with it.

The direction of the skew is named for the long tail, not the bulk. A right-skewed (or positively skewed) batch has most of its numbers bunched on the left with a long tail reaching right; response times behave this way, since most are quick but a few run very long and none can be faster than instant. A left-skewed batch is the mirror image, its long tail reaching left. The quickest tell is to compare the center measures: as a rough guide, the mean gets pulled toward the long tail relative to the median, so a mean noticeably above the median suggests a right lean and a mean below it a left lean. Treat this only as a hint — it points the right way for ordinary humps but can mislead on odd or double-peaked shapes, so confirm it by actually looking at the pile.

Skew matters because it decides which summary is honest. In a strongly skewed batch the mean drifts toward the sparse tail and stops describing a typical member, which is why incomes, durations, and file sizes are so often reported by their median instead. Noticing the lean is what tells you the average has stopped being average.

Turning the pile into a picture

raw values: 3 5 6 6 7 7 7 8 8 11 count into bins 2-4 4-6 6-8 8-10 10-12 bar height = count in that bin
A histogram sorts raw numbers into equal bins and draws a bar per bin whose height is the count that fell into it.

Numbers in a list have no shape you can see. A histogram gives them one. Divide the whole range of values into a row of equal intervals called bins, count how many numbers fall into each bin, and draw a bar whose height is that count. The row of bars is the shape of the distribution made visible: where the bars are tall the numbers crowd, where they are short the numbers thin out.

Read that way, a histogram answers all three questions at once. The center is where the tall bars sit, the spread is how wide the bars extend before they fade to nothing, and a stray bar standing alone far from the rest is an outlier you can point at. A single symmetric hump, a long tail off one side, two separate peaks — each is a distribution you can now describe from the picture instead of the raw numbers.

The one setting that matters is the bin width, and it is a genuine choice, not a detail. Bins too wide smear real structure into a single blunt block; bins too narrow shatter the shape into spiky noise. The same data can look like one clean hump or several ragged peaks depending only on how you sliced it, so the honest move is to try a few widths before trusting any one shape.

The bell and the 68-95-99.7 rule

-3 -2 -1 mean +1 +2 +3 68% within 1 sd 95% within 2 sd 99.7% within 3 sd
In a normal bell, about 68 percent of values fall within one standard deviation of the mean, 95 percent within two, and 99.7 percent within three.

One shape turns up so often it has its own name: the normal distribution, the symmetric bell where numbers crowd near the center and thin out smoothly toward both tails. Measurement errors, heights, and countless natural quantities pile up this way, which is why so much of statistics is built around it. The bell is fixed in form and located entirely by two numbers — its mean, which sets where the peak sits, and its standard deviation, which sets how wide it spreads.

Because the shape is fixed, the standard deviation acts like a ruler laid along the curve, and a reliable pattern called the empirical rule follows. About sixty-eight percent of the values fall within one standard deviation of the mean, about ninety-five percent within two, and about ninety-nine point seven percent within three — the reason it is nicknamed the 68-95-99.7 rule. Nearly everything lives within three standard deviations, so a value four or five out is genuinely rare and worth noticing.

The rule turns a standard deviation from an abstract number into a sense of how surprising a value is. A reading one standard deviation from the mean is ordinary, two is somewhat unusual, three is rare, and beyond that is a strong signal that something is different about that measurement. That instinct only holds when the pile is roughly bell-shaped, so it is worth confirming the shape is normal before leaning on the rule — a strongly skewed batch keeps its own proportions.

Summing up a batch in a few numbers

Any project that measures something ends up with a pile of numbers, and the skill this concept teaches is compressing that pile into a few honest ones. Center, spread, and a glance for outliers are the whole toolkit: a typical value, how much the batch varies around it, and whether any strays are quietly distorting the picture. That summary is what lets you say something true about a batch too large to read.

The capstone projects all lean on it. An agent pipeline produces a step latency on every run, and the median tells you the typical wait while a lone slow retry shows up as an outlier the mean would hide. A driving robot takes a stream of distance readings, and knowing their spread tells signal from a spurious spike. A team tool records how long each ticket took, and reporting the median rather than the mean keeps one marathon ticket from making the whole team look slow. Same three questions, different piles.

The habit to carry away is distrust of a bare average. An average with no spread beside it, and no check for skew or strays, is a number pretending to be a summary. Ask where the center is, how wide the pile spreads, and whether anything sits far from the rest, and you will read a batch of numbers for what it actually says.