Populations and samples — study guide
The concept's fragments, read in order.
The whole you want, the part you get
Ask a simple statistical question — what fraction of voters favor a policy, what a factory's lightbulbs average before they burn out — and the question is secretly about a group far too large to check by hand, one at a time, every member. The group you actually care about is almost never the group you can afford to look at, and nearly everything statistics does starts from that one gap.
The group you care about is the population: every voter, every bulb the factory will ever make, every reading a sensor could produce. What you actually collect is the sample: the handful of voters called, the bulbs pulled off the line and tested, the readings a device happens to log. A sample is not a shrunken copy of the population handed to you for free — it is a separate, much smaller thing that someone had to go collect, and the entire discipline of statistics is a set of tools for saying something honest about the first using only the second, like tasting a spoonful from a pot of soup: you do not need to eat the whole pot to judge it, as long as the spoonful was stirred up and drawn fairly.
That gap between what you want to know and what you can measure is not a flaw waiting to be engineered away. It is the permanent condition under which every real measurement happens, from a single opinion poll to a self-driving vehicle deciding what a stretch of road is like from the last few seconds of sensor data. Naming the two groups precisely — population and sample — is the first move toward getting the gap right instead of being quietly fooled by it.
Population and sample, defined
A population is the complete set of units a question is actually about — every one of them, whether that is three people or three billion. Statisticians write its size as N. A population does not have to be people: it can be every unit a factory will ever produce, every request a server will ever receive, or every possible reading a sensor could take under a given condition. Some populations are so large, or so open-ended, that a full count is not just impractical — it is not even a fixed number waiting to be counted.
A sample is the subset actually collected and measured: the rows that ended up in a dataset, the units a study managed to get its hands on. Statisticians write its size as n, deliberately the lowercase twin of N, so the relationship stays visible in the notation itself — a sample sits inside a population, smaller by definition. A national opinion poll built from 1000 respondents, standing in for a population of adults numbering in the hundreds of millions, is an ordinary example: the sample is tiny next to the population it speaks for, and that ratio is normal, not alarming.
The figure on this page draws that relationship plainly: a full population of units, with a smaller subset picked out and labeled as the sample actually observed. Every quantity computed later in a statistical study — an average, a proportion, a rate — gets computed on the small circled subset first, then used to say something about the much larger set around it.
Parameter and statistic
Every population has some true numeric property — its actual mean, its actual proportion in favor, its actual defect rate — and that number is called a parameter. It is a fixed fact about the population whether or not anyone ever learns its exact value: the true average height of every tree in a forest does not change depending on who is asking, or how many trees get measured.
A statistic is the corresponding number computed from a sample instead — the average height of the trees a survey crew actually measured, the proportion of polled voters who said yes. Statisticians keep the two apart in notation on purpose: population parameters take Greek letters, μ for a population mean and σ for a population standard deviation, while the matching sample statistics take Roman letters, x̄ for a sample mean and s for a sample standard deviation. The split in the alphabet mirrors a split in kind — a parameter is a fixed truth, a statistic is a reading you take, and like any reading it can shift slightly from one taking to the next even though the truth it is estimating has not moved, like a room's actual temperature staying fixed while a thermometer check of it can read a hair higher or lower each time you look, even though the room itself has not changed. The figure on this page marks that fixed value against a computed one, with an arrow showing which is meant to estimate which.
The distinction matters because a parameter and a statistic get used differently in every sentence that follows. A parameter is what a study is trying to learn; a statistic is the tool it uses to guess at it. Treating a computed statistic as if it already were the population's true value is one of the most common ways a sound calculation turns into an overconfident claim.
Why not just measure everyone
The most direct way to learn a population's parameter exactly is a census: measure every single unit, no exceptions, and compute the number directly. Censuses do happen — the United States runs a full population census every 10 years, a cadence written directly into the Constitution — but a full census is the expensive, slow exception, not the normal way statistics gets done.
Most populations resist a census for reasons a once-a-decade headcount does not face. Testing every lightbulb a factory makes for how long it lasts destroys every bulb tested, so a full census would mean selling nothing. Surveying every voter in a country takes long enough that opinions can shift before the count finishes. And plenty of populations are not even finite in a way a census could reach — every reading a sensor could ever produce, every request a server might ever receive, is an open-ended set with no last member left to knock on the door of.
So most statistical work substitutes an ongoing sample for an occasional census. The Census Bureau itself runs one alongside the decennial count: the American Community Survey contacts roughly 3.5 million addresses a year, every year, precisely because a sample can be collected continuously in a way a full census cannot. Sampling is not a shortcut taken because a census would be nice but nobody could be bothered — for most populations, it is the only method that was ever going to work.
What makes a sample representative
A sample is representative when its relevant characteristics mirror the population's — when the mix of whatever matters for the question at hand shows up in the sample in roughly the same proportions it shows up in the population. If a population splits 60 percent one category and 40 percent another, a representative sample of it lands close to that same split, not because anyone forced the numbers to match, but because the sample was drawn in a way that gave every part of the population a fair shot at appearing.
Representativeness is a property of the collection process, not of the number of rows collected. A small sample built well can represent a population far better than a large one built poorly — size describes how much was gathered, representativeness describes whether what was gathered actually looks like the whole. The figure on this page shows exactly that: a population's mix of categories, and a much smaller sample whose mix of the same categories comes out close to matching it, like a scale model of a building: much smaller than the real thing, but built so every part keeps the same proportions as the original.
Exactly how a collection process can fail to be representative, and why that failure resists the obvious fix of simply gathering more, is its own deep subject. What matters here is the target itself: a sample earns the name representative by matching the population's composition, not by size, and not by luck.
The same statistic, computed twice, differs
Draw one sample from a population and compute a statistic from it — a sample mean, a sample proportion. Draw a second, completely separate sample from the exact same population and compute the same statistic again, and it will not come out identical. Different units happened to land in the second draw, so the number built from them differs, usually by a little, sometimes by more. That sample-to-sample wobble has a name: sampling variability.
Sampling variability is not a mistake anyone made. It is the direct, unavoidable consequence of a statistic being built from only part of a population: change which part happened to get collected, and the number built from it moves. Keep drawing fresh samples of the same size from the same population, plot every resulting statistic, and the values scatter around the population's true parameter rather than landing on it exactly — clustered tighter or looser depending on how the population itself varies and how large each sample is.
The figure on this page draws exactly that: several samples pulled from one population, each producing its own statistic, scattered around the fixed value of the parameter they are all estimating. Formally describing the full shape that scatter takes is a subject of its own; the fact worth carrying forward here is simpler — a single statistic from a single sample is one draw from a range of values the same collection process could have produced, never the whole story by itself.
What a larger sample buys you
Grow a sample's size while keeping the same collection process, and the sample-to-sample scatter gets tighter. A statistic built from 1000 observations wobbles less from sample to sample than one built from 10, because more observations average out more of the up-and-down that any one unit contributes. This is a version of the law of large numbers: a statistic computed from independent, identically drawn observations gets closer to the population's true parameter, in probability, as the number of observations grows.
What a larger sample buys is precision — a tighter cluster of possible answers around wherever the collection process happens to be centered. It is not a guarantee of correctness. The law of large numbers' convergence promise depends entirely on that collection process drawing fairly from the population in the first place; if the process itself is skewed toward some part of the population and away from another, collecting more observations tightens the estimate around the skewed answer just as confidently as it would around a correct one. A precise number and a correct number are not the same claim, and sample size alone speaks to only one of them.
So a growing sample size is worth quoting only alongside a claim about how the sample was collected. Precision earned from volume is real and useful — a poll built from 1000 well-drawn respondents genuinely pins down its answer tighter than one built from 50 — but it answers "how tightly clustered," not "clustered around what." Those are two different questions, and only the collection process, never the row count, answers the second.
Where this shows up later
An automated pipeline built around a language model is never evaluated against every possible input it might ever receive — that population is unbounded, and new inputs arrive the moment the system ships. What gets evaluated is a sample: a fixed set of test cases, prompts, or scenarios standing in for the far larger population of everything the pipeline will eventually face. Every number an evaluation report states — a pass rate, an accuracy figure — is a statistic computed on that sample. It estimates a parameter, the pipeline's true real-world performance, that nobody can measure directly, and a fresh batch of test cases would produce a slightly different number even if nothing about the pipeline changed.
A self-driving vehicle's sensors face the same structure in miniature. The population is every driving condition the vehicle could ever meet — every angle of sunlight, every patch of road, every obstacle — and a single test drive samples only a sliver of it. A handful of sensor readings from one afternoon's drive is a sample standing in for that unbounded population, and treating it as anything more than an estimate drawn from whatever conditions happened to occur that day is the same mistake as treating any small sample as the whole truth.
Naming the population, the sample, the parameter, and the statistic precisely, before trusting a single number either system produces, is not academic caution — it is the difference between a report that says what it actually knows and one that quietly claims more than its data can support.