Read this as a study guide instead

Random variables

Attaching numbers to random outcomes so you can do math on them.

The chance process
What the random variable maps each outcome to
How likely each face is
8 outcomes P(X = 2) = 3/8HHH HHT HTH HTTTHH THT TTH TTT0123

The same ideas, as prose

These are the exact fragments the model serves — also available as an ordered study guide.

Turning outcomes into numbers

A random process gives you outcomes, and probability fundamentals already tell you how likely each outcome is. A random variable is the next move: a fixed rule that assigns exactly one number to every outcome, so the messy business of chance becomes something you can compute with. Roll a die and the outcome is a face; treat that face as a random variable and the outcome is a number you can add, compare, or plot.

The rule has to be total and single-valued: every outcome gets a number, and no outcome gets more than one — like a scoreboard that turns each play into a number of points, no matter how different the plays looked on the field. Nothing about the process itself changes when you do this. You are not making the die fairer or the coin less random; you are attaching a consistent numeric label to whatever the process hands you.

That single change is what makes the rest of probability and statistics possible. Once outcomes are numbers, questions like "what number should I expect" or "how spread out are the results" become questions about arithmetic on those numbers — questions this concept sets up but does not answer.

One outcome, one number

one outcome, one number outcome HH HT TH TT count of heads 0 1 2
The four equally likely outcomes of two coin flips, HH, HT, TH, and TT, each connect by exactly one arrow to their count of heads on a number line: HH to 2, HT and TH both to 1, and TT to 0.

Formally, a random variable is a function from the sample space to the real numbers: every outcome the random process could produce goes in, and exactly one real number comes out. Mathematicians write it as a capital letter acting like any other function, taking a whole outcome and returning a single value.

The outcome itself does not have to look like a number at all. Flip a coin and the raw outcome is heads or tails, not 1 or 0 — the random variable is the rule that says heads becomes 1 and tails becomes 0. Flip two coins and the raw outcomes are pairs like heads, tails; a random variable that counts heads maps that pair to 1, and it maps the pair tails, heads to 1 as well. Two different outcomes are allowed to land on the same number — the only rule that can never break is that each outcome lands on exactly one.

That is the whole mechanism, and it is worth seeing before anything else: the mapping is fixed in advance, it covers every outcome the sample space contains, and it never sends one outcome to two different numbers at once.

Countable values vs. a continuum

discrete: a countable list of values example: number of heads in a few flips 0 1 2 3 4 continuous: any value in an unbroken range example: a measured time start end
A discrete random variable's possible values are isolated dots at countable positions on a number line, while a continuous random variable's possible values fill an unbroken segment of the line.

Random variables split into two families by what their possible values look like. A discrete random variable can only take a countable set of values — you could, in principle, list them: 0, 1, 2, and so on, even if the list is long or never ends. The number of heads in a handful of coin flips is discrete; so is the sum showing on a pair of dice.

A continuous random variable takes values from an unbroken range instead — any point along a stretch of the real line is possible, not just a countable list of them. A measured time, a distance, or a voltage reading are the usual examples: between any two possible values there is always another one in between, all the way down.

The distinction matters because the two families are handled with different mathematics from here on. This concept works entirely with the discrete case, where "how likely is each value" has a clean, countable answer; the continuous case needs tools of its own that build on these foundations here.

Spreading probability across the values

A discrete random variable does not just have a list of possible values — each value has its own probability, inherited from however likely the outcomes that map to it were in the first place. Add up the probability of every outcome that lands on a given value, and you get the probability of that value. Do this for every value the random variable can take and you get its full probability mass: one number per value, describing exactly how the total likelihood is split up.

That split cannot be arbitrary. Every outcome in the sample space belongs to exactly one value, and probability fundamentals already guarantee the probabilities of all outcomes add to 1. Carry that fact through the mapping and the probabilities of all the values must add to 1 as well — like slicing one pie among every possible value: the slices can be different sizes, but they always add back up to the whole pie. A value nobody's outcomes reach gets probability 0; a value every outcome reaches gets probability 1; most values land somewhere honest in between.

This probability-per-value assignment has a name — the probability mass function — and it is the complete description of a discrete random variable's behavior. Nothing about the variable is left unaccounted for once you know every value's mass.

"Worked example: the sum of two dice"

sum of two dice: probability out of 36 pairs 1/36 2 2/36 3 3/36 4 4/36 5 5/36 6 6/36 7 5/36 8 4/36 9 3/36 10 2/36 11 1/36 12 sum shown along the bottom the eleven probabilities above add to 1
A bar chart of the sum of two dice: sums 2 through 12 along the horizontal axis, each bar's height equal to that sum's exact probability out of 36 equally likely pairs, shortest at 2 and 12 and tallest at 7, with all eleven probabilities adding to 1.

Roll two fair six-sided dice and the raw outcome is a pair, like 3, 5 — thirty-six equally likely pairs in total, since each die has six faces and the dice do not influence each other. Define a random variable that maps each pair to its sum, and every one of those thirty-six pairs lands on a value between 2 and 12.

The pairs are not spread evenly across the sums, because more than one pair can produce the same total. Only one pair makes 2 (1, 1) and only one makes 12 (6, 6), so those sums each get probability 1/36. Six different pairs make 7 — 1,6, 2,5, 3,4, 4,3, 5,2, and 6,1 — so 7 gets probability 6/36, which reduces to 1/6 and is the single most likely sum. The counts climb from 1 at the ends up to 6 in the middle and back down again, and every one of those eleven probabilities, from 1/36 to 6/36, is exactly its count of pairs divided by 36.

Add all eleven probabilities together — 1/36 eleven ways, weighted by count — and the total is exactly 1, because every one of the thirty-six original pairs was counted in exactly one sum. Nothing is left over and nothing is double-counted; the probability mass function for the sum of two dice accounts for the whole sample space.

"Worked example: heads in three flips"

Flip a fair coin three times and the raw outcome is a sequence like heads, tails, heads — eight equally likely sequences in total, since each flip independently comes up one of two ways and there are three flips (2 times 2 times 2). Written out, they are HHH, HHT, HTH, HTT, THH, THT, TTH, and TTT.

Define a random variable that counts the number of heads in the sequence, and each of those eight sequences maps to a value between 0 and 3. Exactly one sequence has zero heads (TTT) and exactly one has three (HHH), so those values each get probability 1/8. Three sequences have exactly one head (HTT, THT, TTH) and three have exactly two (HHT, HTH, THH), so 1 and 2 each get probability 3/8.

The same check applies as with the dice: 1/8 plus 3/8 plus 3/8 plus 1/8 is exactly 1, because every one of the eight sequences was counted under exactly one value of the count. Same mechanism, same guarantee, a different random process feeding it.

Reading and writing X

Random variables get their own notation, and it is worth learning early because it is used everywhere from here on. A capital letter — usually X — names the random variable itself, the whole rule mapping outcomes to numbers. A lowercase letter, usually the same letter in lowercase (x), names one particular value the random variable could take.

P(X = x) reads as "the probability that the random variable X takes the value x", and it is the probability mass function evaluated at that one value. Swap in a specific number for x — P(X = 7) — and the expression asks for the mass at exactly that value, the kind of number a dice-sum or coin-flip table computes directly.

The convention is only notation, not new mathematics: X is still just the fixed mapping from outcomes to numbers, and P(X = x) is still just "add up the probability of every outcome that maps to x." Writing it this way pays off later, once the same X gets reused inside longer expressions that this concept does not build.

Why attaching numbers pays off

Once an outcome is a number instead of a face, a word, or a sequence, it stops being something you can only describe and starts being something you can compute with. A random variable is what turns "the die came up high" or "the sensor read something unusual" into a value that arithmetic can touch — adding results together, comparing one run against another, feeding a number into whatever comes next.

That is exactly the shape of the readings a lot of real systems produce: a sensor on a moving vehicle reports a number, not a description; a program making a randomized choice branches on a number, not a story about the choice. In every case the underlying process is still random, but the random variable is what makes the randomness usable by the rest of the system.

This concept stops at the mapping and its probability mass — it does not say what a typical value looks like or how spread out the values are. Those questions, and the arithmetic that answers them, are exactly what attaching a number to every outcome was for.