Prefer to click through the interactive model?
Correlation and covariance — study guide
The same fragments the interactive model serves, read in order. One source, two views.
Two numbers, moving together
Variance already gives a precise number for how far a single random variable wanders from its own average. Covariance and its scaled cousin the correlation coefficient extend that same idea to a pair of variables at once, asking not how far either one strays alone but whether their straying lines up. A day's temperature and that day's electricity use tend to rise on the same days; a household's income and its grocery spending tend to sit at similar rungs of the same ladder. Both are cases of two numbers moving together, and "moving together" turns out to be a shape precise enough to measure.
Two variables can each have a perfectly ordinary spread on their own and still be locked together, stepping up and down in near lockstep like like two dancers whose steps track each other beat for beat, one stepping up exactly as the other does, without either one leading the other, or they can be just as spread out individually and show no consistent relationship between them at all. A plain description of each variable by itself — its own mean, its own variance — cannot tell those two situations apart, because neither one says anything about the other variable.
None of what follows says anything about which variable, if either, is driving the other; that is a separate question with its own separate machinery. The goal here is narrower and comes first regardless: given two paired columns of numbers, compute one honest number for how tightly they move together, know exactly what that number can and cannot promise, and know precisely where it breaks.
"Covariance: the signed average of paired deviations"
Covariance measures co-movement the same way variance measures spread: by averaging a deviation from the mean, except now there are two variables and two means. For paired variables X and Y with means mu_x and mu_y, the population covariance is Cov(X,Y) = E[(X - mu_x)(Y - mu_y)] — for every paired outcome, subtract each variable's own mean from its own value, multiply the two deviations together, and average that product over every pair, weighted by how likely each pair is. When the underlying population isn't directly available and a sample of n paired observations stands in for it, the sample covariance instead divides the summed products by n - 1 rather than n, the same correction variance already needs once its own mean has to be estimated from the sample rather than known outright.
The sign of that averaged product is the whole story of direction. A pair whose X deviation and Y deviation share a sign — both above their means together, or both below together — contributes a positive product; a pair where one deviation is positive and the other negative contributes a negative product. Covariance is whichever kind of pair wins out on average: positive when high-X pairs tend to go with high-Y, negative when high-X pairs tend to go with low-Y.
One property of the number itself is easy to miss and matters later: covariance carries units, and not the original ones. A covariance between a variable measured in dollars and one measured in hours comes out in dollar-hours, a compound unit with no direct meaning of its own — a warning sign that covariance's raw size, unlike its sign, is not yet the whole story.
Computing a covariance by hand
Five paired measurements are enough to carry the formula through by hand. Take hours studied for a quiz, X = 1, 2, 3, 4, 5, alongside that student's score out of ten on five different quizzes, Y = 2, 3, 5, 4, 6. The mean of X is 3; the mean of Y is 4.
Subtracting each mean from its own column gives the deviations dx = -2, -1, 0, 1, 2 and dy = -2, -1, 1, 0, 2, one pair per quiz. Multiplying each pair of deviations together gives the products 4, 1, 0, 0, 4, and adding those five products gives 9. Dividing by the number of pairs, 5, lands on a population covariance of 1.8; had this been a sample standing in for a larger population, the same sum divided by n - 1 = 4 instead would give 2.25.
Three of the five points fall in a quadrant where both deviations share a sign — both negative or both positive — contributing positive products of 4, 1, and 4; the other two sit exactly on one mean line, one with a zero X deviation and the other with a zero Y deviation, so their products vanish and neither helps nor hurts the sum. Every contribution that survives is positive, which is exactly why the total lands positive: more hours studied went with higher quiz scores across this small dataset, and the covariance's sign says precisely that much, nothing more.
The same relationship, a different-looking number
Rescale one of the worked example's variables and covariance moves in lockstep with the rescaling, even though nothing about how the two variables actually relate has changed. Reporting that same quiz score out of 1000 instead of out of 10 — multiplying every Y value by 100 — multiplies every deviation from the mean by 100 too, and since covariance sums products of deviations, the whole covariance scales by 100: the population covariance of 1.8 becomes 180, computed from the identical five students in the identical order like reporting a distance as a bigger number simply because it was measured in inches instead of miles, with nothing about the actual distance having changed at all.
Nothing about which students studied more or scored higher changed between those two versions of the dataset — only the ruler the second variable happens to be measured on. A covariance of 180 is not a stronger relationship than a covariance of 1.8; it is the same relationship, reported in different units.
That leaves covariance unable to answer a question it looks like it should: given two different pairs of variables, which pair moves together more tightly? Comparing their raw covariances says as much about which pair happened to be measured in bigger units as it does about which pair actually clusters more closely around a line. A number that strips the units back out is what the correlation coefficient builds.
"r: covariance measured in standard deviations"
The fix for covariance's scale problem is to divide it by the very thing that is scaling it: each variable's own standard deviation. The correlation coefficient, usually written r, is defined as r = Cov(X,Y) / (sigma_x * sigma_y), covariance divided by the product of the two standard deviations. Dividing a quantity by something measured in its own units cancels those units exactly, so r carries no units at all — it is a pure number, comparable across any two variables regardless of what either one is measured in.
The worked five-student example makes this concrete. Both X and Y have a population variance of 2, so both standard deviations equal the square root of 2, about 1.41. Dividing the population covariance of 1.8 by that product, 2, gives r = 0.9. Switching to the sample version changes every ingredient — the covariance becomes 2.25, each variance becomes 2.5, each standard deviation about 1.58 — but the n - 1 correction appears in the covariance's numerator and in both standard deviations' denominators alike, and it cancels out of the ratio completely: 2.25 / (1.58 * 1.58) still comes out to 0.9.
That cancellation is not a coincidence of this dataset; it holds for any data, which is why r never needs a population-versus-sample footnote the way covariance and variance do. A rescaled version of the same dataset confirms the point from the other direction: multiplying Y by 100 multiplies its standard deviation by 100 too, so the new covariance of 180 divided by the new, hundred-times-larger standard deviation product lands back on exactly 0.9. Reporting the quiz out of 1000 instead of 10 changed covariance by a factor of 100 and changed r not at all.
Why r can never leave [-1, 1]
The bound on r is not an empirical pattern noticed after the fact; it falls straight out of a property variance already guarantees. Standardize both variables first — subtract each one's mean and divide by its own standard deviation, producing Zx and Zy, each with mean 0 and variance exactly 1 — and r turns out to equal Cov(Zx, Zy) exactly, since standardizing is itself the division r performs. Variance can never be negative, because it averages squared numbers, so Var(Zx - Zy) and Var(Zx + Zy) are both guaranteed to be 0 or greater.
Expanding those two variances the same way a variance of any sum or difference expands gives Var(Zx - Zy) = Var(Zx) + Var(Zy) - 2*Cov(Zx,Zy) = 2 - 2r and Var(Zx + Zy) = Var(Zx) + Var(Zy) + 2*Cov(Zx,Zy) = 2 + 2r. Neither can be negative, so 2 - 2r >= 0 forces r <= 1, and 2 + 2r >= 0 forces r >= -1. There is no dataset, however contrived, that can push r outside [-1, 1] — the bound is a consequence of variance's own non-negativity, not a limit anyone had to impose from outside.
The same standardized view also explains why r survives a rescaled variable unchanged, and why scale dependence was covariance's problem alone. Multiplying a variable by a positive constant before standardizing changes nothing about its standardized version, since the standard deviation absorbs exactly that constant back out; multiplying by a negative constant flips every standardized value's sign, which flips r's sign along with it. r is blind to how a variable happens to be scaled, and blind in a precise, provable way — not just in this worked example, but for any linear rescaling of either variable.
What r=0 does not tell you
Take five paired values, X = -2, -1, 0, 1, 2 and Y set to exactly X squared: Y = 4, 1, 0, 1, 4. Every single Y value is completely determined by its X value — there is no noise anywhere in this dataset, and knowing X tells you Y exactly, always. If a relationship this tight doesn't drive r toward one extreme or the other, nothing should.
Working the covariance formula through anyway: the mean of X is 0, the mean of Y is 2, and the five deviation products come out to -4, 1, 0, -1, 4. Those five numbers sum to exactly 0, which makes the covariance 0 and, since dividing zero by anything nonzero still gives zero, r = 0 as well — the same number r reports for two variables with no relationship whatsoever.
The parabola is not a loophole or an edge case; it is r doing exactly what its formula was built to do. r measures how well a single straight line fits the data, and only that. A relationship that curves — rising then falling, or falling then rising, the way this one does on either side of X = 0 — cancels out in a sum of straight-line deviation products, because the rising half and the falling half of the curve pull r in opposite directions and average away, even though the underlying pattern is perfectly real and perfectly predictable.
That gap between r = 0 and no relationship at all is the single most consequential thing to get right about this number. A small r rules out a straight-line relationship; it never rules out a relationship. Whether some other, non-linear pattern is hiding behind an r near zero is a question r alone cannot answer, and answering it takes a closer look at the actual shape of the data, not a bigger dataset or a fancier formula.
Where this number does real work
A formal, unit-free number for how tightly two variables move together turns out to be one of the most reused tools in any system built on real data. A pipeline that strings several automated steps together often has a cheap, fast signal and an expensive, slow one meant to measure roughly the same thing; checking whether the two actually correlate, and how tightly, is the first honest test of whether the cheap signal is worth trusting on its own. The same question shows up wherever two streams of paired numbers exist side by side and someone needs to know whether they carry the same information or two different ones.
A machine built to sense its own surroundings — reading several sensors at once, each noisy in its own way — runs into the identical question in a different shape: do two sensor readings drift together because they are measuring related things, or do they wander independently, each carrying information the other one lacks? Covariance and r are the exact tools for putting a number on that, before any decision gets made about how to combine the two readings or trust one over the other.
None of that decision-making happens here. What this concept hands off is narrower and more durable: a way to compute, from any two paired columns of numbers, one honest, comparable, unit-free measure of how tightly they move together, plus a precise account of what that measure promises and where its promise runs out.