Read this as a study guide instead

Linear regression

Fitting the best straight line through data, and reading its coefficients.

y = 1.3 + 0.9x SSR = 1.9 R^2 = 0.81residuals -0.2, -0.1, 1, -0.9, 0.2 add to 0

The same ideas, as prose

These are the exact fragments the model serves — also available as an ordered study guide.

Fitting a line through a cloud of points

Two variables that move together, the way correlation and covariance measure, invite an obvious next question: what is the actual rule connecting them? Linear regression answers it by fitting a single straight line through a scatter of data points — not any line that looks about right, but the one specific line that fits best by a precise, agreed-upon rule.

That line comes with two numbers worth reading on their own: a slope, which says how much the second variable changes for every one-unit change in the first, and an intercept, which says where the line sits when the first variable is zero. Together they turn a cloud of dots into an equation, one that can be read off, quoted, and used to predict a value at an x the data never actually showed.

The rest of this concept works through exactly what "best" means, computes it by hand on one small dataset, and marks the places where trusting the line too far goes wrong.

What makes one line better than another

two lines through the same five points worse guess least-squares line x=1 x=2 x=3 x=4 x=5 squared residuals sum: 21.90 (worse guess) vs 1.90 (least squares)
Five data points with a worse candidate line and long dashed residual segments to it, next to the least-squares line with much shorter solid residual segments; the worse line's sum of squared residuals is 21.90 while the least-squares line's is 1.90.

Any line drawn through the general drift of a set of points can look plausible at a glance, but two lines can both look roughly right and still miss the data by very different amounts. Something has to define, precisely, which one line the fitting procedure actually chooses.

For every point, measure the vertical distance between the observed value and the value the line would have predicted at that point's x — that gap is called a residual, and it can run positive or negative depending on whether the point sits above or below the line. Least-squares regression squares each residual before adding them up, then picks the one line, out of every line that could possibly be drawn, that makes that sum as small as possible.

Squaring the residuals rather than just summing the raw distances matters for two reasons: a positive miss and a negative miss of the same size no longer cancel each other into a false zero, and a miss twice as large counts four times as much, so the line ends up compromising most for the points it would otherwise miss badly — like a tug-of-war where every point pulls the line toward itself, and squaring the distance means the biggest misses pull hardest of all.

On the small dataset used throughout this concept, a deliberately worse guess at the line — one with the same slope but an intercept pushed two units too high — racks up a sum of squared residuals of 21.90, while the least-squares line brings that same sum down to 1.90, the smallest total any line through these five points can achieve.

Computing the slope and intercept

the fitted line: y = 1.3 + 0.9 times x 0 run = 1 rise = 0.9 intercept: (0, 1.3) slope = 0.9
The least-squares line, 1.3 plus 0.9 times x, for the five-point dataset, with a slope triangle showing a rise of 0.9 for a run of 1, and the intercept marked where the line crosses x=0 at y=1.3.

The least-squares slope has a closed-form answer, not a search: for every point, multiply its distance from the mean of x by its distance from the mean of y, add those products together across every point, and divide by the sum of the squared distances from the mean of x alone. That ratio is exactly the same number as the covariance of x and y divided by the variance of x, or equivalently the correlation coefficient r scaled by the ratio of the two standard deviations.

Five points make the arithmetic concrete: x values 1, 2, 3, 4, 5 paired with y values 2, 3, 5, 4, 6. The mean of x is 3 and the mean of y is 4. Each point's two distances from those means multiply out to 4, 1, 0, 0, 4, summing to 9, and the squared distances from the mean of x alone sum to 10. The slope is therefore 9 divided by 10, exactly 0.9.

The intercept follows from a fact true of every least-squares fit that includes one: the fitted line always passes through the point where both variables sit at their own mean. Solving for the intercept means asking what number, added to the slope times the mean of x, gives back the mean of y: intercept + 0.9 times 3 = 4, so the intercept is 4 - 2.7, exactly 1.3. The fitted line for this dataset is y = 1.3 + 0.9 times x, predicting 2.2, 3.1, 4.0, 4.9, 5.8 at x = 1 through 5.

Why the line always passes through the average point

the line pivots through the mean point mean x = 3 mean y = 4
The least-squares line passing exactly through the point where x equals 3 and y equals 4, the mean of the five data points, marked at the center of the scatter with dashed guide lines to each axis.

Every least-squares line, fit to any dataset with an intercept term, passes exactly through one specific point: the mean of x paired with the mean of y. For the five-point dataset used throughout this concept, that point is x = 3, y = 4 — and the fitted line, 1.3 + 0.9 times x, lands on it exactly: 1.3 + 0.9 times 3 works out to 1.3 + 2.7, which is 4.

The reason is built into how the intercept gets computed, rather than being a coincidence discovered after the fact: the intercept is defined as the mean of y minus the slope times the mean of x, precisely so that plugging the mean of x back into the line returns the mean of y. No matter which slope the least-squares criterion settles on, the intercept formula forces the line to pivot through that one mean point — like a seesaw that always balances at the exact average position of the weights sitting on it, however those weights are arranged.

This gives every fitted line a fixed anchor regardless of how the surrounding points scatter: two datasets that share the same mean x and mean y but disagree everywhere else will still produce lines that cross at that identical point, only tilted differently around it.

Residuals: what the line misses

residuals: observed minus predicted, one per point -0.2 -0.1 +1.0 -0.9 +0.2 -0.2 - 0.1 + 1.0 - 0.9 + 0.2 = 0.0
Vertical residual segments from each of the five data points to the least-squares line, labeled with their signed values -0.2, -0.1, +1.0, -0.9, and +0.2, which sum to exactly zero.

A residual is the gap between what actually happened and what the fitted line predicted: the observed y minus the predicted y, computed separately at every point. A positive residual means the actual point sits above the line; a negative one means it sits below.

For the five-point dataset, the residuals at x = 1 through 5 are -0.2, -0.1, 1.0, -0.9, 0.2 — the point at x = 3 overshoots the line by a full unit, while the point at x = 1 misses by only two tenths.

Add those five residuals together and the total is exactly 0.0, and that is not a property of this particular dataset — it holds for every least-squares fit that includes an intercept term. The same calculation that sets the intercept also guarantees the positive and negative misses balance out exactly, so a fitted line can never systematically run high or low across the data it was fit to; whatever bias exists gets absorbed into the intercept itself.

R-squared: how much the line explains

Every point in a dataset varies from the mean of y by some amount, and squaring and summing those deviations gives a single number for how much y varies overall. R-squared asks how much of that total the fitted line accounts for: it is the share of the total squared deviation that the line explains, leaving the rest to the residuals.

For the five-point dataset, the total squared deviation of y from its own mean is 10, and the residuals left over after fitting the line sum, when squared, to 1.9. R-squared is 1 minus 1.9 divided by 10, exactly 0.81 — the line accounts for 81 percent of how much y varies, and the remaining 19 percent is scatter no straight line through these five points can explain.

For a single-variable fit like this one, R-squared is not an independent number to compute separately: it is exactly the square of the correlation coefficient r. This dataset's correlation is 0.9, and 0.9 squared is 0.81, matching the direct calculation exactly — a strong r and a high R-squared are the same fact stated two ways, though that exact equivalence stops holding once more than one predictor variable enters the picture.

Where the line stops being trustworthy

A fitted line is only justified over the range of x values the data actually covered. The five-point dataset used throughout this concept only says something about x between 1 and 5; plugging in x = 50 produces a number the formula is happy to compute but that nothing in the data actually supports, since no point anywhere near that value was ever observed.

The term regression itself comes from a distinct but easily confused idea: in the late 1800s, Francis Galton noticed that unusually extreme measurements — very tall parents, an especially high test score — tend to be followed by more ordinary ones, a pattern he called regression toward the mean, like a run of unusually hot days that is almost always followed by a run of more ordinary ones, not because the heat is being corrected but because extremes rarely repeat themselves. That statistical drift toward average happens on its own in any pair of imperfectly correlated measurements — a repeated measurement of the same quantity included — and is a separate phenomenon from the line-fitting technique that happens to share its name.

A least-squares line also assumes the underlying relationship is genuinely straight and that no single point is distorting the fit; a curved relationship or one wild outlier can still produce a line with a deceptively confident-looking equation. Judging whether those conditions actually hold is its own careful exercise, one this concept only flags rather than works through.

Where a fitted line does real work

A handful of runs through an evaluation pipeline rarely line up perfectly, but plotting a score against a changing setting and fitting a line through the results reads the trend the same way it gets read here: a slope worth quoting, and a line to eyeball rather than staring at scattered numbers one run at a time.

A raw sensor on a small rolling chassis almost never reports the unit anyone actually wants — a distance sensor hands back a voltage, not centimeters. Two calibration readings, one near each end of the sensor's working range, are enough points to fit the same kind of line used throughout this concept and convert every future raw reading into the physical unit it was always meant to represent.

Both uses lean on nothing more than the slope, the intercept, and the residuals covered here — turning that same fitted line into something closer to a working model, one built on the assumptions that make a fit trustworthy rather than just computable, is its own concept.