The same ideas, as prose
These are the exact fragments the model serves — also available as an
ordered study guide.
What the memory hierarchy is
Your computer does not keep data in one place. It keeps it in several places at once, arranged in layers from tiny-and-fast to huge-and-slow: registers inside the CPU itself, then one or more levels of cache, then RAM, then disk. Each layer down holds far more data than the one above it — and takes far longer to answer.
The arrangement is a deal with physics. Fast storage must sit close to the CPU and stay small; big storage sits farther away and answers slowly. So the hardware keeps the data it expects you to need next in the small fast layers, and reaches down to the big slow ones only when it has to. It works like doing research at a desk: the page in your hands is a register, the papers spread on the desk are cache, the shelf across the room is RAM, and the library across town is disk.
The model on this page steps through real reads. Some are answered near the top, almost instantly. Others miss, fall through layer after layer, and the running counter shows the cost piling up. That accumulating wait — not any single layer — is the thing to understand.
Registers — the CPU's hands
Registers are storage built into the CPU itself: a few dozen named slots, each holding one machine word — 8 bytes on a 64-bit machine. Reading a register takes a fraction of a nanosecond, because it happens inside the same clock cycle as the instruction using it. There is no trip anywhere; the data is already in the CPU's hands.
Everything the CPU actually computes with passes through registers. An add instruction adds two registers; a comparison compares them. Data anywhere else — cache, RAM, disk — must be brought into a register before the CPU can touch it, which is why compilers work hard to keep the busiest values living in registers for as long as possible.
Registers are the top of the hierarchy and its extreme case: the fastest storage a computer has, and by far the smallest — a few hundred bytes in total, against gigabytes of RAM.
Cache levels and cache lines
Between registers and RAM sit the caches: small blocks of very fast memory on the CPU chip. L1 is the closest and smallest — typically 32 to 64 kilobytes per core, answering in about a nanosecond. L2 is larger, often around a megabyte, and a few times slower. Many chips add a still larger shared L3. Each level trades a little speed for a lot more room.
Caches do not store single bytes. They store cache lines — fixed blocks, typically 64 bytes. When a cache pulls data up from below, it pulls the entire line containing it, not just the byte that was asked for. That detail looks wasteful and is anything but: it is what makes the next read of a neighboring value nearly free.
A cache holds recently used lines and quietly evicts the ones used least. Nothing about it is visible in your code — the hardware manages it — but whether your data happens to be in a line the cache kept is often the difference between a fast program and a slow one.
RAM — main memory
RAM is the computer's working memory: the gigabytes that hold your running programs and their data. It lives on separate chips outside the CPU, reached over a memory bus, and a read takes roughly 100 nanoseconds — about a hundred times slower than L1 cache. In CPU terms that is a long pause: a modern core could have executed hundreds of instructions in the time one RAM read takes.
When RAM answers a cache miss, it does not send back a lone byte — it sends the whole 64-byte cache line, which the caches keep. One slow trip pays for many fast ones afterward.
RAM is also where the hierarchy's memory stops being permanent. It is volatile: cut the power and its contents are gone. Anything that must survive belongs one level further down, on disk.
Disk — where data survives
Disk is the bottom of the hierarchy: terabytes of storage that keep their contents with the power off. The price is time. A solid-state drive answers a read in around 100,000 nanoseconds — a thousand times slower than RAM — and a spinning hard drive is another hundred times slower still, because a physical arm has to move to the data.
Programs do not read the disk directly the way they read memory. The operating system moves data between disk and RAM in chunks called pages, and your code only ever touches the copy in RAM. The first touch of data that has not been loaded yet is what forces the full trip: every cache misses, RAM comes up empty, and the operating system goes to disk.
That full trip is the most expensive ordinary thing a program does. A single disk read costs as much time as hundreds of thousands of cache hits — which is why so much of computing, from databases to operating systems, is really the art of not touching disk.
The gaps are the lesson
The hierarchy's levels are not evenly spaced — the slowdowns come in leaps. A register read costs about half a nanosecond, L1 about 1, L2 about 4, RAM about 100, and an SSD about 100,000. Each leap down multiplies the wait, and the leaps get bigger the further you fall.
Numbers that small are hard to feel, so scale them up: if a register read took one second, an L1 hit would take two seconds, a RAM read over three minutes, and a single SSD read more than two days. That is the terrain your programs cross on every read — and why where data lives matters more than almost anything else about it.
The leaps also decide what is worth optimizing. Shaving cache hits saves nanoseconds; avoiding one trip to disk saves the equivalent of days. When the model's counter jumps by a hundred thousand at the final step, that jump is the entire argument of this concept in one number.
Cache hit, cache miss
Every read poses the same question at each level: do you have this data? If the level has it, that is a hit — the read is answered at that level's speed and the trip ends. If not, that is a miss: the request falls through to the next level down and asks again, and the costs add up along the way. It is like reaching for a paper that is already on your desk versus discovering it is not there and walking across town to the library to fetch it.
A miss is not one cost but a chain of them. Read a value no cache has seen and you pay the L1 check that fails, the L2 check that fails, and then the full RAM trip — around 105 nanoseconds where a hit would have cost one. The failed checks themselves are cheap; the trip they force is not.
The fraction of reads answered by a cache is its hit rate, and small changes in it move a program's speed dramatically, because each miss costs a hundred or more hits' worth of time. A cache that hits 99 times in 100 spends most of its time — the actual nanoseconds — on the one read that missed.
Locality — why caches work at all
A cache is a bet, and the bet has a name: locality. Temporal locality is the observation that data you used a moment ago, you will probably use again soon. Spatial locality is the observation that data sitting next to what you just used is probably next in line. Real programs — loops, arrays, structs — follow both patterns so reliably that a tiny cache can answer the vast majority of reads.
The 64-byte cache line is spatial locality made physical. When the model's miss finally reaches RAM, RAM sends back the entire line, so the missed value's neighbors ride along for free. That is why the very next step — a read of the neighboring value — hits in L1: one expensive trip quietly prepaid for the cheap reads after it.
This is also the rare part of the hierarchy your code can influence. A loop that walks an array in order moves with the cache lines and hits almost every time; code that chases pointers to scattered places defeats the bet and misses constantly. Same data, same machine — very different speed.
Why the hierarchy runs the show
A modern CPU core can start several instructions every nanosecond. Set that against the hierarchy's numbers and the imbalance is stark: one RAM read costs a few hundred instructions' worth of waiting, and one disk read costs hundreds of thousands. The instruction cycle you learned about earlier only runs at full speed while its data is close; most of a slow program's time is not computation at all, but a fast CPU standing idle while data makes the trip up.
That reframes what "slow" usually means. When a program crawls, the first question is rarely whether the CPU is fast enough — it is where the data is coming from. Waiting on RAM, on disk, or on anything even further away dwarfs the cost of the arithmetic in between.
The same shape repeats at every scale of computing: the hierarchy continues past disk to the network, where fetching from another machine costs more than disk by the same kind of leap. Keep data close to where the work happens — it is one of the few ideas that explains performance from a single function all the way up to a global system.