Read this as a study guide instead

Sampling and bias

Who made it into the dataset, who didn't, and why more data won't fix it.

valuedraw #trueframe

Code with its results

A notebook-style walk through the idea — every output shown is the real result of the code above it.

A biased sample that a bigger N can't save

Run a sampling experiment and watch the sample mean lock onto a skewed frame's value instead of the true population value, no matter how large N gets.

Real output Every Out block below was produced by running the code above it. You can copy these cells into your own python3 and run them top to bottom to see the same numbers.

An agent pipeline logs the latency of each run, but only runs that finished. The runs that crashed never made it into the log, so the logged runs are a biased frame drawn from all runs. Only statistics and a seeded random are used.

In [1]
import statistics
import random
all_runs = list(range(1, 101))
logged = [x for x in all_runs if x <= 60]
print("true mean latency (all runs) =", statistics.mean(all_runs))
print("logged-runs mean latency =", statistics.mean(logged))
Out [1]
true mean latency (all runs) = 50.5
logged-runs mean latency = 30.5

Draw samples from the logged runs and let N grow. The mean latency settles near the logged frame's mean, which undercounts the slow, failed runs.

In [2]
rng = random.Random(0)
for n in [10, 100, 1000, 10000]:
    sample = rng.choices(logged, k=n)
    print("N =", n, "logged sample mean =", round(statistics.mean(sample), 2))
Out [2]
N = 10 logged sample mean = 32.7
N = 100 logged sample mean = 35.32
N = 1000 logged sample mean = 30.12
N = 10000 logged sample mean = 30.56

For contrast, draw the same growing N from all runs, crashed ones included. The mean tightens toward the true latency.

In [3]
rng = random.Random(0)
for n in [10, 100, 1000, 10000]:
    sample = rng.choices(all_runs, k=n)
    print("N =", n, "all-runs sample mean =", round(statistics.mean(sample), 2))
Out [3]
N = 10 all-runs sample mean = 54
N = 100 all-runs sample mean = 58.58
N = 1000 all-runs sample mean = 49.86
N = 10000 all-runs sample mean = 50.61

Separate the wobble from the gap. At a large N the all-runs error has all but vanished, while the logged-only error stays stuck below zero: the reported latency is confidently too low.

In [4]
rng = random.Random(0)
logged_big = rng.choices(logged, k=10000)
all_big = rng.choices(all_runs, k=10000)
true_latency = statistics.mean(all_runs)
logged_err = statistics.mean(logged_big) - true_latency
all_err = statistics.mean(all_big) - true_latency
print("logged error at N=10000 =", round(logged_err, 2))
print("all-runs error at N=10000 =", round(all_err, 2))
Out [4]
logged error at N=10000 = -19.96
all-runs error at N=10000 = 0.02

The same ideas, as prose

These are the exact fragments the model serves — also available as an ordered study guide.

Who made it into the data

A dataset is not the world. It is whoever the collection process happened to let in, and the gap between those two things is where sampling bias lives.

Sampling bias is a systematic mismatch between the group you care about and the group your data actually describes. It is not random noise that averages out; it is a lean baked into how the rows were gathered, present before a single calculation runs. If the people — or planes, or program runs, or sensor readings — that got recorded differ in some consistent way from the ones that didn't, every number you compute inherits that lean.

The trap is that a biased dataset looks exactly like a healthy one. It has real rows, real values, a believable center and spread. Nothing on the surface announces who is missing. The whole skill is learning to ask, before you trust any figure, who the collection quietly left out — and why more rows will not bring them back.

Population, frame, sample

population sampling frame sample left out by the frame (the bias gap)
The sample sits inside the sampling frame, which sits inside the population; the slice of the population the frame never covers is the bias gap, and no draw from the frame can reach it.

Three groups sit nested inside one another, and naming them tells you exactly where bias gets in.

The population is everyone or everything you actually want to know about: every voter, every flight, every driving condition the car will ever meet. The sampling frame is the narrower group you can actually reach — the list you draw from, the mechanism that decides who is even eligible to be recorded. The sample is what you finally collect out of that frame. Bias is the silent slice the frame leaves out of the population, not bad luck in the draw but a gap built into the list itself, like polling voters out of a phone book in an era when poorer households had no phone: a fair draw from the book still cannot reach anyone who was never listed in it.

This is the distinction that trips people. A perfectly fair, perfectly random draw from a crooked frame is still crooked, because randomness only protects you against picking unevenly from the list you have. It does nothing about who never made the list. If your frame omits a group, no sampling procedure downstream can sample them, because as far as the data is concerned they do not exist.

So the first question about any dataset is not how large it is or how it was sampled, but what the frame was — the concrete list or mechanism the rows were drawn from, and who that mechanism could never have reached.

The classic ways a sample goes crooked

Sampling bias arrives in a handful of recognizable shapes, and each one bends the frame at a different moment.

Selection bias is the broad case: the frame itself over- or under-represents part of the population, so the eligible pool is already skewed before anyone is chosen. Convenience sampling is selection bias by laziness — you take whoever is easy to reach, the people walking past or the users who already signed up, and quietly treat them as the whole. Non-response bias, also called self-selection, strikes after you have a frame: you invite everyone, but the people who choose to answer differ systematically from the people who ignore you, so the responders speak for a group that is not the population.

Survivorship bias is the sneakiest, because the missing entries removed themselves. You measure only the things that made it to the end — the funds still trading, the machines still running, the planes that came back — and the failures are simply absent, dragging every average toward the survivors.

The names matter less than the move underneath all of them: in every case some group had a lower chance of ending up in your data than another, and that difference, not the size of what remains, is the bias.

The planes that came back

damage seen on planes that returned cockpit armor here: hits here downed the missing planes
Returning bombers show damage on the wings and fuselage; because planes hit in the engines and cockpit did not return to be measured, those unmarked areas are where armor is actually needed.

The clearest picture of survivorship bias comes from a wartime story about armor, and it is worth telling carefully, because the popular version has grown well past the record.

During the Second World War, the Statistical Research Group at Columbia studied the damage on Allied bombers returning from missions, looking for where to add armor. The tempting move was to reinforce wherever the returning planes showed the most bullet holes. Abraham Wald, a statistician in the group, is credited with the crucial reframing: the data described only the planes that came back. As the often-told account has it, he reasoned that armor belonged where the survivors were unmarked — the engines and cockpit — because a plane hit there was a plane that never returned to be measured.

Treat the dramatic details as legend rather than documentation. Wald's own reports never use the phrase survivorship bias, name no specific armor zones, and the idea of armoring the least-damaged areas was already understood by wartime militaries; the memorable retelling accreted later. What survives scrutiny is the logic, and it is exactly right: the sample was survivors only, so the places it showed no damage were precisely the places that were fatal.

That is the shape to carry away. When a dataset contains only the things that made it, the gaps in the damage are not proof of safety — they are the evidence that got shot down.

Why more data won't fix it

growing sample size N true value biased frame value bias: fixed gap
As sample size grows both estimates get more precise, but the unbiased one closes on the true value while the biased one closes on the wrong value; the gap between them, the bias, never shrinks.

Here is the load-bearing lesson, the one that separates bias from every other data problem: a bigger sample does not cure it.

Two different things can be wrong with an estimate. Variance is random wobble — draw a fresh sample and the number jumps around, sometimes high, sometimes low. Bias is a systematic offset — the number is consistently off in the same direction, no matter how the draw goes. Collecting more data attacks only the first. Averaging over more rows cancels the random wobble, so variance shrinks toward zero as the sample grows. But every new row still comes from the same skewed frame, so it carries the same lean, and that offset never averages away, like a bathroom scale that reads five pounds heavy: weigh yourself a thousand times and average the results, and the answer is still five pounds too high.

Picture the estimate tightening as the sample grows. If the frame is sound, it tightens onto the true value. If the frame is skewed, it tightens just as confidently onto the wrong value — and now you have a precise, stable, well-supported number that is still wrong. More data can even make things worse, by dressing a biased answer in the false authority of a huge sample.

This is why "we will just collect more" is the wrong reflex for a biased dataset. The fix for bias is never volume; it is fixing the frame, reaching the group the collection was missing. A crooked ruler measured a thousand times reports its error to more decimal places, not fewer.

Two million wrong answers

In 1936 one poll collected millions of responses and still called the American presidential election spectacularly wrong. It is the case that retires the idea that size alone saves you.

The magazine Literary Digest mailed out around ten million mock ballots and tallied about 2.4 million returns — a sample so large that a modern poll, which surveys a few thousand, looks like rounding error beside it. It predicted Alf Landon would beat Franklin Roosevelt with roughly 57 percent of the vote. Roosevelt won in a landslide, taking about 62 percent and all but eight of the electoral votes. The poll had missed by nearly twenty points, one of the largest errors in polling history.

The failure was in the frame, not the count. The Digest built its mailing lists from telephone directories, automobile registrations, its own subscribers, and club rosters — a comfortably-off slice of a country still in the Depression, and one that leaned Republican. Then non-response compounded it: only about a quarter of the ballots came back, and the people who bothered to return them were not a cross-section of even that skewed list. Millions of responses drawn from the wrong frame produced a very confident wrong answer.

The counterpoint is the punchline. George Gallup called the race correctly for Roosevelt with a sample of only about fifty thousand — hundreds of times smaller, but drawn to resemble the actual electorate. As the story is usually told he even forecast that the Digest would get it wrong, though how firmly he predicted that in advance is disputed. A small, well-framed sample beat a giant crooked one, and it has been the textbook lesson ever since.

Interrogating the frame

You cannot fix a bias you never looked for, so the habit worth building is interrogating the frame before you trust the number.

The questions are always the same. What was the population I actually care about? What list or mechanism did these rows really come from, and who could it never have reached? Who was invited but chose not to appear? What got recorded only because it survived long enough to be recorded? None of these are answered by staring harder at the values you have; they are answered by reconstructing how the data came to exist. And when a group turns out to be missing, the honest response is to go get it or to state the limit plainly — not to average the survivors more carefully.

This lands hardest the moment you collect your own data. Pull the logs from an agent pipeline and you are studying survivors: the runs that crashed may never have been recorded at all, so an average latency computed over the logged runs quietly excludes the slowest, failing ones. Gather driving data from your own car and the frame is simply whatever you happened to drive through — collect only in daylight on a clean test track and the dataset will confidently describe a world with no rain, no dusk, and no traffic.

The dataset in front of you is always a sample of some frame. Knowing whose story it can and cannot tell is the difference between a measurement and a mirage.