Prefer to click through the interactive model?

Variance — study guide

The same fragments the interactive model serves, read in order. One source, two views.

How far a random variable strays from its average

A random variable's expected value tells you where it lands on average, but nothing about how confidently it lands there. Two random variables can share the exact same expected value and still behave completely differently — one clustering tightly around that number, the other swinging wildly above and below it. Variance is the formal measure that tells the two apart: it captures how spread out a random variable's possible outcomes are around its own expected value.

Nothing about variance depends on which particular outcome shows up on a given trial. It is a property of the whole random variable — every possible outcome weighted by how likely it is — computed once and true of every future draw. A variance of zero means the random variable never actually varies: every outcome equals the expected value. Any larger number means real spread, and how much spread is exactly what a precise number needs to state.

Getting from "spread out" to an actual number takes one specific move: measuring each outcome's distance from the expected value, then squaring that distance before averaging it. Why squaring, rather than something simpler, turns out to explain most of what makes variance useful.

Why plain deviations cancel out

signed deviations from the mean — they cancel to zero mean = 3.5 1 -2.5 2 -1.5 3 -0.5 4 +0.5 5 +1.5 6 +2.5 -2.5 -1.5 -0.5 +0.5 +1.5 +2.5 = 0
Six signed deviations of a fair die's outcomes from its mean of 3.5, three negative bars below the mean line and three positive bars above it, matching lengths that cancel exactly to zero when added.

The most natural way to measure spread might seem to be averaging how far each outcome sits from the expected value, sign and all. That plan fails immediately: for any random variable, the average of the raw signed deviations from its own expected value is always exactly zero. Outcomes above the mean and outcomes below it pull in opposite directions, and by the very definition of the mean, those pulls cancel perfectly like a tug-of-war that always ends in a dead heat: the pulls above the mean and the pulls below it are always exactly matched, so the rope never actually moves — a tug-of-war that always ends in a dead heat, no matter how strong either side's outcomes actually are.

Squaring each deviation before averaging fixes the cancellation, because a squared number is never negative regardless of which direction the original deviation pointed. It has a second effect too: it does not treat every unit of distance the same. A deviation twice as large contributes four times as much once squared, so a single big miss counts for far more than several small ones added up like scoring a miss so that being twice as far off the target costs four times as many points, not just twice as many. That property is not an accident of the arithmetic: it is why variance is sensitive to outliers in a way a plain average distance never is, and it is also what lets the variances of independent random variables add together directly, a convenience no absolute-value measure offers.

Squaring is also the choice that keeps the mean itself as the natural center of a random variable. The expected value is exactly the number that minimizes the average squared deviation; measuring distance with an absolute value instead would make the median, not the mean, the minimizing center. Since expected value is already the quantity a random variable is built around, squared deviations are the measure that stays consistent with it.

"The definition: expected squared deviation"

Variance gives the intuition of "expected squared deviation" a precise formula. For a discrete random variable X with expected value E[X] (written mu), the variance is Var(X) = E[(X - mu)^2] — take every possible outcome, subtract the mean, square the result, and average those squared distances weighted by how likely each outcome actually is. Written out over the random variable's own possible values, that is Var(X) = sum of (x - mu)^2 times P(X = x), one term per outcome.

Every piece of that formula does real work. The (x - mu) term is the raw deviation for one particular outcome; squaring it removes the sign and weights larger misses more heavily; multiplying by P(X = x) makes sure outcomes that almost never happen barely move the total, while outcomes that happen often pull their full weight. Summing across every outcome folds the whole random variable into a single number.

That single number carries squared units of whatever the random variable measures — variance of a random variable measured in dollars comes out in dollars squared, which is exactly why it is rarely reported on its own.

Working it out for a fair die

squared deviations from the mean, one bar per face 6.25 face 1 2.25 face 2 0.25 face 3 0.25 face 4 2.25 face 5 6.25 face 6 each weighted 1/6, sum = 35/12 ≈ 2.92
Six bars, one per fair-die face 1 through 6, each bar's height exactly proportional to its squared deviation from the mean 3.5, each weighted equally at probability 1 in 6, summing to the variance 35 over 12.

A single roll of a fair six-sided die is small enough to compute by hand and concrete enough to make the formula stick. Each face 1 through 6 comes up with probability 1/6, and the expected value works out to 3.5 — the average of the six faces, exactly midway between 3 and 4 even though 3.5 never itself appears on a die.

Each face's deviation from 3.5 is -2.5, -1.5, -0.5, 0.5, 1.5, and 2.5 for faces 1 through 6. Squaring each of those gives 6.25, 2.25, 0.25, 0.25, 2.25, and 6.25, and weighting every one by the shared probability 1/6 before adding them together gives the variance: Var(X) = (6.25 + 2.25 + 0.25 + 0.25 + 2.25 + 6.25) / 6 = 35/12, a little under 3, or about 2.92.

That 35/12 is not a rounded estimate; it is the exact variance of a fair die roll, and any other route through the same random variable's numbers should land on that identical fraction.

"A faster route: E[X^2] minus E[X] squared"

Squaring and averaging the deviations directly always gives the right answer, but there is a faster route that avoids finding the mean twice: Var(X) = E[X^2] - (E[X])^2, the expected value of the squared random variable minus the square of its expected value. The two formulas are not separate facts to memorize; the second falls straight out of expanding the first.

Expanding (X - mu)^2 gives X^2 - 2*mu*X + mu^2. Taking the expected value of each term separately — expectation distributes over a sum — turns that into E[X^2] - 2*mu*E[X] + mu^2. Since mu is itself E[X], the middle term becomes 2*(E[X])^2 and the last term is (E[X])^2, so -2*(E[X])^2 + (E[X])^2 collapses to -(E[X])^2, leaving exactly E[X^2] - (E[X])^2.

The fair die confirms it. Its expected value squared is 3.5^2 = 12.25, and E[X^2], the average of each face squared weighted by 1/6, works out to 91/6, about 15.17. Subtracting gives 91/6 - 12.25 = 35/12, the identical variance the direct definition produced — the shortcut is not an approximation, it is the same number reached by less arithmetic.

"Standard deviation: back to the original units"

standard deviation: an equal distance on each side of the mean mean mean - sigma mean + sigma sigma sigma same units as the outcomes — variance itself is in squared units
A number line with the mean marked at its center and a standard-deviation bracket of equal length on each side, in the same units as the outcomes, unlike variance's squared units.

Variance answers "how spread out," but its units make it awkward to interpret directly, since squaring the deviations also squares the units they were measured in. A random variable measured in dollars has a variance measured in dollars squared; a variable measured in seconds has a variance in seconds squared — a number with no direct counterpart in the world the random variable describes.

Standard deviation fixes that by undoing the squaring: it is defined as the square root of the variance, sigma = sqrt(Var(X)). Taking the square root brings the units back to the original scale, so a random variable measured in dollars gets a standard deviation measured in dollars, one that can sit next to the expected value and mean the same kind of thing. For the fair die, the standard deviation is the square root of 35/12, a little under 1.71.

Standard deviation is best read as a typical distance an outcome sits from the expected value like the typical walk from home to wherever an errand actually ends up: not the longest possible trip, not one every errand makes, but a single distance that sums up how far outcomes tend to wander — not the largest possible distance, and not a distance every outcome actually achieves, but a single number that summarizes how far the random variable's outcomes tend to wander. A small standard deviation means outcomes cluster close to the mean; a large one means they routinely land far from it.

How shifting and scaling change variance

original — two deviation bars around the mean shift by b — bars unchanged, only the mean moves slid by b scale by a — bars stretch by a, squared deviations scale by a^2 each bar longer by the same factor a
Three paired number lines: the original two deviation bars, an unchanged copy slid sideways after adding a constant so variance is unchanged, and a copy with both bars stretched by the same factor after multiplying by a constant so variance scales by that factor squared.

Variance behaves in two very specific ways when a random variable is shifted or scaled by constants, and both follow directly from the squared-deviation definition. For a random variable X and constants a and b, Var(aX + b) = a^2 * Var(X) — the additive constant b disappears entirely, while the multiplicative constant a reappears squared.

Shifting makes sense once the deviations are examined: adding b to every possible outcome also shifts the expected value by exactly b, so every outcome's distance from the, also shifted, mean is completely unchanged. Nothing about the spread of the random variable moved, only its location did, and variance measures spread, not location.

Scaling works differently because multiplying every outcome by a also multiplies every deviation from the mean by a. Since variance sums squared deviations, multiplying each deviation by a multiplies each squared deviation by a^2, and that factor carries straight through the sum untouched. A common mistake is assuming the scale factor carries through unsquared, the way it would for the expected value itself: Var(aX + b) is not, in general, a * Var(X) + b, however tempting the parallel to E[aX + b] = a * E[X] + b looks.

Why variance is worth tracking

Variance is the quantity underneath almost every later argument about risk, noise, and uncertainty, because "how spread out" is usually the more consequential question once "what is the average" has already been answered. A self-driving system reasoning about a noisy sensor reading, a pipeline of agents reasoning about how reliable one step's output is, and a small team estimating how long a project will actually take are all, underneath the specifics, asking the same question variance answers: not just the expected outcome, but how far a real outcome might land from it.

Two estimates can share an identical expected value and still deserve completely different treatment once their variance is known — the estimate with near-zero variance can be trusted almost as stated, while the one with high variance needs a plan for what happens when reality lands far from the average. That distinction is invisible to the expected value alone, and it is exactly what variance was built to expose.

Variance itself, together with the standard deviation built from it, is also the foundation the next layer of ideas is built on: it is the quantity a whole probability distribution's shape gets measured against, the number later summary statistics report about real data, and the piece two random variables need before their relationship to each other can even be asked about. None of that gets explained here — variance is the piece being handed off, not the argument being made with it.