Prefer to click through the interactive model?
Probability distributions — study guide
The same fragments the interactive model serves, read in order. One source, two views.
A complete rulebook for every outcome
An expected value is one number, a summary of a random variable's long-run average, and a variance is another, describing how far outcomes typically stray from it. A distribution is more ambitious than either: it is the complete map from every possible value a random variable can take to the exact chance of the variable landing there, all in one object like a rulebook that lists every possible outcome in advance and states exactly how likely each one is, before anything actually happens. Knowing a variable's distribution means knowing everything probability can tell you about it, including both of those summary numbers as figures you can compute from the whole picture rather than facts you need separately.
Most distributions worth knowing are not one-off inventions. A small number of named shapes — the binomial, the uniform, the normal, and a handful of others — recur across wildly different situations, from counting coin flips to modeling measurement error, because the same underlying mechanism produces them again and again. Each family is defined by a formula plus a small set of parameters: fix the parameters and the entire distribution, every likelihood it assigns, is pinned down.
Learning a distribution family, then, is learning two things at once: the shape the formula produces, and what its parameters mean in the world the model describes. The rest of this concept works through three of the most common families and how to recognize which one fits a given question.
Mass at points, density over intervals
A random variable is discrete when its possible values can be listed one at a time — a count of successes, a number rolled on a die — and its distribution is described by a probability mass function, or PMF. The PMF assigns an actual probability to each individual value, a genuine number between 0 and 1, and those numbers across every possible value add up to exactly 1. Asking for the probability of one specific outcome is always a sensible question for a discrete variable, and the PMF is the direct answer.
A continuous random variable can take any value in a range — a measurement, a duration — and there are too many possible values to list, let alone assign each one a nonzero probability. Its distribution is described instead by a probability density function, or PDF, a curve over the range whose height at a point is not itself a probability. Probability for a continuous variable comes from the area under that curve over an interval, and the total area under the whole curve is exactly 1, the continuous counterpart to the PMF's sum.
The consequence is easy to state and easy to forget: a continuous variable has probability exactly 0 of landing on any single exact value, because a single point has no width and so no area under the curve. What has probability is always a range — between 2 and 3, say — never one number picked out exactly. Every distribution from here on is one or the other, mass at points or density over intervals, and within this concept the two never mix within a single variable.
When every outcome is equally likely
The uniform distribution is the flattest possible model: every outcome in its range is exactly as likely as every other, with no outcome favored. A fair six-sided die is the simplest example — each face from 1 to 6 carries probability 1/6, no more and no less, because nothing about the die favors one face over another. That flatness is itself the whole model; there is no shape to describe beyond constant.
The continuous version works the same way over an interval instead of a list of faces. For a uniform distribution on the interval from a to b, the density is the constant 1/(b-a) everywhere inside the interval and 0 everywhere outside it — a flat rectangle sitting on top of the range, with the height chosen so the rectangle's area comes out to exactly 1. A wider interval means a shorter rectangle; a narrower one means a taller one, since the total area can never change.
Uniformity is an assumption, and like any assumption it needs to be earned rather than granted by default. It fits a die roll or a spin of a well-balanced wheel, where symmetry genuinely rules out favoring any outcome. It does not fit a quantity with a natural center of gravity — heights, reaction times, measurement error — where values near the middle really are more common than values at the extremes, and a flat model would misdescribe them badly.
Counting successes in a fixed number of trials
The binomial distribution models a specific, very common situation: repeating the same yes/no trial a fixed number of times and counting how many times it comes up a success. Each trial succeeds with the same probability p and is independent of every other trial like flipping the same weighted coin a fixed number of times, each flip blind to how any of the others landed — flipping a coin n times and counting heads is the textbook case, but the same pattern covers anything with a fixed success probability repeated a fixed number of times. Two parameters, the trial count n and the success probability p, are the entire model.
The probability of exactly k successes out of n trials is C(n,k) · p^k · (1-p)^(n-k): the chance of one particular pattern of k successes and n-k failures, p^k · (1-p)^(n-k), multiplied by C(n,k), the count of how many different orderings of successes and failures produce that same total. That combinatorial factor is doing real work — it is why counts near n times p end up far more likely than the extremes, even though every specific ordering with the same number of successes is exactly as probable as any other on its own.
A concrete case makes the shape visible: five independent coin flips, each a fair p = 0.5. The probability of exactly k heads for k running 0 through 5 comes out to 1/32, 5/32, 10/32, 10/32, 5/32, and 1/32 — symmetric, peaked in the middle, and summing to 32/32, which is 1, as any complete PMF must.
Reading a binomial's parameters back out
Every distribution has a mean and a variance, computable in general the same way any random variable's are: weight each possible value by its probability and add. The binomial distribution is unusual in how little work that takes, because its own two parameters already contain the answer. The mean of a binomial with n trials and success probability p is simply n · p, and the variance is n · p · (1 - p) — no summing required once n and p are known.
The five-flip example makes the shortcut concrete: with n = 5 and p = 0.5, the mean is 5 × 0.5 = 2.5 and the variance is 5 × 0.5 × 0.5 = 1.25. A mean of 2.5 heads out of five flips is exactly what symmetry already suggested — the peak of the distribution sits right at the middle, between 2 and 3 successes, because 2.5 itself is not a possible outcome.
The pattern behind both formulas is the same independence that built the distribution in the first place: each trial contributes p to the mean and p(1-p) to the variance on its own, and because the trials do not influence each other, those n individual contributions simply add. Push p toward 0 or 1 and the variance shrinks toward 0 as well — a lopsided coin gives a far more predictable count of heads than a fair one, and the formula n · p · (1 - p) says so directly, since p(1-p) is largest exactly at p = 0.5.
The bell curve and its two parameters
The normal distribution is a continuous model shaped like a single symmetric bell: a peak at the center, thinning out evenly in both directions, never touching zero but getting close fast. Its density is given by the formula f(x) = (1 / (σ√(2π))) · exp(-(x - μ)^2 / (2σ^2)), which looks formidable but is controlled by exactly two parameters. μ (mu) is the mean, fixing where the peak sits, and σ (sigma) is the standard deviation, fixing how wide the bell spreads before it thins out.
Nothing else about the curve is free to vary. Two normal distributions with the same μ and σ are the identical curve, full stop, and changing either parameter moves or reshapes the whole bell in a predictable way: a larger μ slides the peak sideways, a larger σ flattens and widens it without moving the center.
The shape shows up constantly because of how it tends to arise: whenever a quantity is the sum of many small, independent influences pulling roughly evenly in different directions, the result tends to cluster into this exact bell like many small nudges pushing something in random directions at once: they tend to cancel out near the center and thin out evenly the further out they would have to pile up in one direction. Measurement error, human heights within a population, and test scores across a large group all take this shape for that reason, which is why the normal distribution earns a name of its own rather than being just one pattern among many.
The 68-95-99.7 rule
For any normal distribution, no matter its μ or σ, the same fixed shares of probability sit within a given number of standard deviations of the mean. About 68% of the distribution falls within one standard deviation of μ on either side, about 95% falls within two, and about 99.7% falls within three — the 68-95-99.7 rule, sometimes called the empirical rule, and it is a standard approximation worth memorizing in that rounded form.
Because the shares depend only on how many standard deviations away a value is, not on the specific μ or σ, the rule applies identically to a normal distribution with a small spread and one with a huge spread — both put about 68% of their probability within one σ of center, just measured on very different scales. That is what makes the standard deviation a genuinely useful unit for a normal distribution: it converts "how many raw units away" into "how unusual," on a scale that means the same thing no matter which normal distribution is in play.
The flip side is where the rule gets useful in practice: a value more than three standard deviations from the mean captures only about 0.3% of the distribution combined across both tails, which is why an observation that far out reads as a genuine outlier rather than ordinary variation. The three bands nest inside each other — the 68% band sits entirely inside the 95% band, which sits entirely inside the 99.7% band — so moving one standard deviation further out always trades a little more coverage for a lot more width.
Picking the distribution that fits the question
Picking a distribution family starts with one question: is the quantity a count or a measurement? A count — how many trials succeeded, how many events occurred — is discrete, and a fixed number of independent yes/no trials with a constant success probability points straight at the binomial. A measurement — a length, a duration, an amount — is continuous, and the question shifts from listing outcomes to describing a density.
Among continuous quantities, the next question is whether every value in some range is equally plausible or whether values cluster toward a center. Flat plausibility across a bounded range, with a genuine symmetry argument behind it, is the uniform distribution's signature. Clustering around a center, produced by many small independent influences adding up, is the normal distribution's signature instead — and the two look nothing alike once drawn, a flat rectangle against a peaked bell.
None of this is guesswork forever. Real data can be checked against a proposed distribution, and a scenario's own structure — trials versus measurements, symmetry versus central tendency — usually narrows the choice before any checking starts. What a chosen distribution buys is a compact model: instead of tracking every observation individually, later reasoning about a quantity's uncertainty can work directly from a distribution's own parameters and shape.