Prefer to click through the interactive model?
Parallel processing — study guide
The same fragments the interactive model serves, read in order. One source, two views.
Two shapes of processor
A CPU and a GPU are built from the same raw material — transistors etched into silicon — but they spend it on opposite bets. A CPU pours its budget into a few big, clever cores, each able to tear through one complicated task quickly. A GPU spreads its budget across thousands of small, simple cores that are individually unremarkable and collectively overwhelming.
Neither shape is better; they answer different questions. The CPU asks "how fast can I finish this one thing?" The GPU asks "how many things can I finish at once?" It works like a sports car against a fleet of buses — the car gets a couple of people there fast, while the fleet moves far more people per hour even though each trip is slow, and which one you want depends entirely on the trip.
This concept is about that shape difference: why a processor meant for raw parallel scale looks so unlike the one running your operating system, what that shape is brilliant at, and what it is hopeless at. The whole story follows from a single design choice — many simple workers instead of a few powerful ones.
Finish one fast, or finish many
There are two ways to be fast, and they pull against each other. Latency is how long one task takes from start to finish. Throughput is how much work gets done per second across everything in flight. A CPU is built to minimize latency — get this one thing done now. A GPU is built to maximize throughput — get an enormous pile of things done, even if any single one waits its turn.
That is why the two chips look different inside. A CPU spends most of its transistors on caches and control logic that keep a single task moving without stalling. A GPU spends most of its transistors on raw arithmetic units and keeps thousands of tasks in flight at once, so that whenever some of them are waiting, others always have work to do — it hides the waiting behind sheer volume. It works like one express checkout against twenty ordinary ones — the express lane clears your single basket quickest, but the twenty lanes clear far more shoppers per hour.
The payoff is a clean rule of thumb. If you have one urgent task, you want low latency and the CPU wins. If you have a mountain of similar tasks, you want high throughput and the GPU wins by a landslide.
One instruction, many data
The trick that makes thousands of cores manageable is to stop giving each one its own instructions. Instead, one instruction is broadcast to many cores at once, and each applies that same operation to its own piece of data. This is called SIMD — single instruction, multiple data: the same operation performed on many data points simultaneously.
It works like an aerobics instructor calling one move that a whole room performs at once, each person on their own mat but everyone on the same beat. Because every core is doing the identical step at the same moment, the hardware needs only one instruction decoder feeding a whole row of arithmetic units, which is exactly what frees up the transistor budget to have so many of them. The cost is rigidity: the cores move in lockstep, all on the same beat.
That single design decision is the source of both the GPU's power and its limits. When your work really is the same operation over a huge array of data, SIMD is a superpower. When your work needs each element to do something different, marching everyone in lockstep stops helping.
Many simple cores beat a few clever ones
The core counts are not close. A modern consumer CPU has a few to a few dozen cores — roughly 4 to 24 on mainstream desktops and laptops, and up to around 64 on high-end parts, as of the mid-2020s. A modern GPU has thousands, from a couple thousand on modest cards to over ten thousand on high-end ones; a high-end board like the GeForce RTX 4090 lists 16384 of its cores.
The GPU can afford that many only because each one is stripped down. A CPU core carries deep caches and elaborate control logic so a single thread never stalls; a GPU dedicates most of its transistors to arithmetic instead, spending almost nothing per core on cleverness. Simple and numerous beats complex and few — but only for the right work.
That qualifier is the whole game. Thousands of cores are worthless if the job cannot be cut into thousands of independent identical pieces; then most of them sit idle while a handful do the work, and the CPU's few fast cores would have finished sooner. The GPU's advantage is real only when the work fills its width.
What GPUs are good for
The work a GPU loves has a name: embarrassingly parallel, meaning it takes little or no effort to split into a pile of independent tasks that need not talk to each other. When the same operation applies to millions of separate data elements, a GPU is in its element.
Graphics is the original case and the reason GPUs exist: every pixel on the screen can be shaded by the same little program, all at once. The same shape turns up far from graphics. Multiplying large matrices is thousands of identical independent multiply-add sums, which is why linear algebra runs so well on a GPU. Training neural networks — a later concept — is mostly that same matrix math, which is why the whole field runs on GPUs.
The common thread is uniform work over independent data: same operation, different element, no waiting on the neighbors. When a problem has that shape, handing it to thousands of cores turns hours into seconds.
What they're bad at, and how they team up
Give a GPU the wrong shape of work and its thousands of cores turn into a liability. A task that must happen in strict order cannot be spread across lanes at all — the GPU's width goes to waste while the CPU's fast single core would have finished it sooner. Heavy branching is nearly as bad: when the cores run in lockstep and one needs to take a different path than its neighbors, the hardware has to walk both paths in turn, so divergent code serializes the very parallelism it was meant to exploit.
So the two chips are not rivals; they are a team with a clear division of labor. The CPU runs the program, handles the branchy, sequential, decision-heavy work, and when it hits a big uniform parallel chunk it offloads that chunk to the GPU, then collects the results when the GPU is done. On a machine with a separate graphics card, that hand-off travels across a bus called PCIe.
The rule that falls out is worth keeping: reach for the GPU when the work is wide and uniform, stay on the CPU when it is twisty and sequential, and expect most real programs to need both.
Reading a GPU spec
Once the shape makes sense, a GPU spec sheet reads like plain language. The core count — thousands, versus the handful on a CPU — tells you how wide the chip is, how many identical pieces of work it can run at once. The memory figure tells you how big a batch of data can sit on the card at a time, close to those cores. Both numbers describe the same bet: enormous parallel width, fed by lots of nearby data.
This is why a self-driving RC car leans on a GPU. Reading a camera frame and deciding what is road, obstacle, or sign is a neural network doing the same arithmetic across every patch of the image — wide, uniform, and constant. That is exactly the work a handful of CPU cores would choke on and thousands of GPU cores devour, frame after frame, fast enough to steer.
The takeaway outlasts any one project: match the chip to the shape of the work. When you see a job that is the same operation over a mountain of independent data, you now know why the answer is a GPU — and why, when the job is one twisty sequence, it is not.