Prefer to click through the interactive model?

Correlation and causation — study guide

The same fragments the interactive model serves, read in order. One source, two views.

Moving together isn't the same as making it happen

Two columns of numbers can rise and fall in near-perfect step, and that fact alone tells you almost nothing about why. Correlation is the pattern of moving together; causation is one thing making another happen. They are different claims, and the data that shows the first does not, by itself, earn you the second.

The gap matters because the leap across it is so tempting. When sales climb on the days a new banner is up, the banner feels responsible. When people who take a supplement live longer, the supplement feels like the reason. The co-movement is real in both cases; the causal story is a separate thing you have to argue for, not read off the chart.

This concept is mostly a discipline of restraint: seeing a strong relationship, naming it honestly as co-movement, and then stopping short of the conclusion the numbers do not license. The rest of these fragments are the specific ways that conclusion goes wrong, and the one kind of evidence that actually supports it.

Co-movement, and how a scatter shows it

positive: they rise together negative: one falls as the other rises
A scatter that tilts up shows positive correlation; a scatter that tilts down shows negative correlation.

Correlation starts with paired measurements: for each thing you observe, you record two numbers at once — a person's height and their weight, a day's temperature and its electricity use. Plot each pair as a point, one measurement along the bottom and the other up the side, and you get a scatter: a cloud of dots whose shape is the relationship.

When the cloud tilts upward, with points on the right sitting higher than points on the left, the two measurements tend to rise together, and the correlation is positive. When it tilts downward, one tends to fall as the other rises, and the correlation is negative. When the cloud is a shapeless blob with no tilt, there is little or no linear correlation: knowing one number tells you nothing useful about the other.

That tilt is the whole of what correlation measures — a tendency for two numbers to move in a consistent direction together. It is a statement about the shape of a cloud of dots, and nothing more. It says which way they move together, not that either one is steering.

r, one number for the co-movement

-1 0 +1 tight, tilts down no tilt tight, tilts up sign is direction; distance from 0 is tightness
The correlation coefficient r runs from minus one to plus one: the sign is the direction, and the magnitude toward one is how tightly the points hug a line.

A scatter's tilt can be summarized in a single number, the correlation coefficient, written r. It runs from -1 to +1. The sign is the direction: positive r for a cloud that tilts up, negative r for one that tilts down. The magnitude — how close r sits to 1 in either direction — is the tightness: how nearly the points hug a straight line instead of scattering loosely around it.

At the extremes, r equals +1 or -1 only when every point falls exactly on one sloping line. An r near 0 means no linear tilt at all. So an r of 0.9 is a strong positive relationship, -0.8 a strong negative one, and 0.05 essentially nothing. Python's standard library will compute it for you: statistics.correlation(x, y) returns exactly this number for two lists of paired values.

The one trap built into r is the word linear. It measures how well a straight line fits, and only that. A relationship that is real but curved — one that rises, peaks, and falls — can have an r near zero even though the two numbers are tightly linked, because no single straight line captures a hill. A small r rules out a straight-line relationship, not every relationship.

The conclusion you don't get to draw

Here is the conclusion a correlation does not give you: that one of the two things causes the other. A strong r says the numbers move together. It is silent on why, and the why has several possible shapes that the co-movement cannot tell apart.

The reflex to fill in causation is old and strong like the old belief that the rooster's crow makes the sun rise, since the crow comes first every single morning. Reliable timing and repetition feel like proof, but they are just the co-movement again, dressed in a story. Before you can say one thing causes another, you have to rule out the other explanations that would produce the very same correlation.

There are three, each worth ruling out: a hidden third thing driving both, a pure coincidence, and the arrow simply pointing the other way. Until those are ruled out, a correlation is a question, not an answer.

The lurking variable behind both

hidden cause: summer heat ice-cream sales drownings causes causes they correlate but neither causes the other
A confounder drives both measured variables, so they correlate even though neither one causes the other.

The first explanation to rule out is a confounder: a third variable that drives both of the things you measured, so they move together without either one touching the other. The classic pairing is ice-cream sales and drownings, which rise and fall in near-lockstep across the year like ice-cream sales and drownings both climbing every summer, not because one drives the other but because the heat drives both. Ice cream does not cause drowning; hot summer weather independently drives people to buy cones and to swim.

Draw it and the shape is unmistakable: an arrow from the hidden cause to each of the two measured things, and no arrow between them. The heat points at ice-cream sales and points, separately, at swimming. The correlation you see between the two is a real shadow of a cause neither of them is. The hidden driver is the lurking variable, and it is doing all the work.

Confounders are what makes observational data so treacherous. The variable doing the real work is often one you did not record, or did not think to, and it can manufacture a strong, stable correlation between two things that have nothing to do with each other. The question worth keeping on hand is: what else could be moving both of these at once?

Coincidence, and the arrow pointing the wrong way

X Y X Y X Y assumed: X causes Y reverse: Y causes X coincidence: no link
The same correlation between X and Y is compatible with X causing Y, Y causing X, or coincidence with no causal link.

Two more explanations can hide behind a correlation. The first is pure coincidence — a spurious correlation, where two unrelated things happen to trace the same wandering line for a while with no shared cause at all. Comb through enough pairs of series and some will match by chance. Tyler Vigen's Spurious Correlations project collects these: across one decade, US spending on science and technology tracked suicides by hanging with a correlation of about 0.99, a figure near the very top of r's range and plainly meaningless.

The second is reverse causation: the relationship is causal, but the arrow runs opposite to the one you assumed. You notice that towns with more police have more crime and conclude police cause crime; far likelier, crime drives towns to hire police. Same correlation, arrow reversed. And because r is symmetric — statistics.correlation(x, y) returns exactly the same value as statistics.correlation(y, x) — the number itself can never tell you which way the arrow points.

So a single correlation is compatible with at least four different worlds: one thing causing the other, the other causing the first, a confounder causing both, or plain coincidence. The coefficient looks identical in all four. Telling them apart takes evidence the scatter plot does not contain.

What it takes to earn a causal claim

random split one group treatment: gets it control: does not compare outcome randomizing balances the confounders across both groups
Random assignment splits one group into a treatment and a control group, balancing confounders so an outcome difference is attributable to the treatment.

If observing more data cannot settle causation, what can? Intervening. Instead of watching two things vary on their own, you step in and change one of them yourself, then see whether the other responds — while holding everything else as steady as you can.

The tool that does this cleanly is the randomized controlled experiment. Take a group, split it at random into a treatment group that gets the intervention and a control group that does not, and compare the outcomes. Randomness is the crucial ingredient: assigning members by coin flip balances every other variable — the ones you know about and the ones you never thought of — evenly across both groups. Whatever difference shows up afterward is left attributable to the one thing you deliberately changed. A well-run randomized controlled trial is considered the gold standard for establishing causation for exactly this reason.

This is why a confident causal claim rests on an experiment, not on a bigger pile of observations. More observational data buys you a more precise correlation, and a more precise correlation is still just co-movement. The randomization is what breaks the confounders that observation alone can never rule out.

When you'll feel this

You will meet this concept the moment you have two columns of your own numbers and a hunch about them. Log an rc_car's runs and you will have throttle, wheel speed, battery voltage, and motor temperature all recorded together, and pairs of them will move in tight step. Throttle and wheel speed rise together — and there the causal story happens to be true. Motor temperature and tire wear rise together too, and there it is not: longer runs simply drive both.

The habit worth building is to separate the two claims every time. First state what the data actually shows: these two numbers correlate, this strongly, in this direction. Only then, as a separate and harder question, ask whether one causes the other — and notice that answering it honestly usually needs an experiment you have not run, not another lap of logging.

That restraint is the whole payoff. It will not tell you which correlations are causal, but it reliably stops you from acting on the ones that are not — which is where the expensive mistakes live.