Read this as a study guide instead

Parallel processing

Why GPUs are shaped differently and what that's good for.

Lanes
Work type
one op → many lanes (lockstep)OP

Check your own machine

These commands only read information — they change nothing. Run the block for your OS, then use the table to read your own numbers against the ideas above.

Windows

Your GPU and its memory PowerShell
Get-CimInstance Win32_VideoController | Select-Object Name, AdapterRAM, DriverVersion

What you should see One row per graphics adapter: Name is the GPU model, AdapterRAM is its memory in bytes, DriverVersion is the installed driver.

Note AdapterRAM is a 32-bit field, so it under-reports any GPU with more than 4 GB of memory. Read the true figure in Settings, System, Display, Advanced display. Windows does not expose the GPU core count.

macOS

Your GPU, its cores and memory zsh
system_profiler SPDisplaysDataType

What you should see The Chipset Model names the GPU. On Apple Silicon a Total Number of Cores line reports the GPU core count; on Macs with discrete or Intel graphics a VRAM line reports the memory instead.

Linux

List your GPU(s) bash
lspci | grep -E 'VGA|3D'

What you should see One line per graphics device: the vendor and model of every GPU the system sees, integrated and discrete.

GPU detail (NVIDIA cards) bash
nvidia-smi

What you should see A table with the card model, its total memory, and current utilization.

Note nvidia-smi ships only with the NVIDIA driver, so it works only on NVIDIA cards. On AMD or Intel graphics use the vendor's own tool (for example radeontop or intel_gpu_top) instead.

Where each idea shows up in your output
ConceptWindowsmacOSLinux
GPU modelNameChipset Modelthe VGA/3D line from lspci
GPU memoryAdapterRAM (under-reports above 4 GB)VRAM (Total)Memory from nvidia-smi (NVIDIA)
Core / compute-unit countnot exposed by WMITotal Number of Cores (Apple Silicon)nvidia-smi or a vendor tool

The same ideas, as prose

These are the exact fragments the model serves — also available as an ordered study guide.

Two shapes of processor

A CPU and a GPU are built from the same raw material — transistors etched into silicon — but they spend it on opposite bets. A CPU pours its budget into a few big, clever cores, each able to tear through one complicated task quickly. A GPU spreads its budget across thousands of small, simple cores that are individually unremarkable and collectively overwhelming.

Neither shape is better; they answer different questions. The CPU asks "how fast can I finish this one thing?" The GPU asks "how many things can I finish at once?" It works like a sports car against a fleet of buses — the car gets a couple of people there fast, while the fleet moves far more people per hour even though each trip is slow, and which one you want depends entirely on the trip.

This concept is about that shape difference: why a processor meant for raw parallel scale looks so unlike the one running your operating system, what that shape is brilliant at, and what it is hopeless at. The whole story follows from a single design choice — many simple workers instead of a few powerful ones.

Finish one fast, or finish many

same time span — two different wins CPU 1 fast lane first result — soonest (low latency) GPU many lanes 5 of thousands of lanes — far more finished per span (high throughput) time
Over one time span, a CPU's single fast lane returns its first result soonest, while a GPU's many lanes finish far more tasks in total.

There are two ways to be fast, and they pull against each other. Latency is how long one task takes from start to finish. Throughput is how much work gets done per second across everything in flight. A CPU is built to minimize latency — get this one thing done now. A GPU is built to maximize throughput — get an enormous pile of things done, even if any single one waits its turn.

That is why the two chips look different inside. A CPU spends most of its transistors on caches and control logic that keep a single task moving without stalling. A GPU spends most of its transistors on raw arithmetic units and keeps thousands of tasks in flight at once, so that whenever some of them are waiting, others always have work to do — it hides the waiting behind sheer volume. It works like one express checkout against twenty ordinary ones — the express lane clears your single basket quickest, but the twenty lanes clear far more shoppers per hour.

The payoff is a clean rule of thumb. If you have one urgent task, you want low latency and the CPU wins. If you have a mountain of similar tasks, you want high throughput and the GPU wins by a landslide.

One instruction, many data

one instruction, many data — every lane runs the same step at once same operation broadcast to all lanes data 1 apply op result 1 data 2 result 2 data 3 result 3 data 4 result 4
One instruction is broadcast to many lanes at once, and each lane applies that same operation to its own data element.

The trick that makes thousands of cores manageable is to stop giving each one its own instructions. Instead, one instruction is broadcast to many cores at once, and each applies that same operation to its own piece of data. This is called SIMD — single instruction, multiple data: the same operation performed on many data points simultaneously.

It works like an aerobics instructor calling one move that a whole room performs at once, each person on their own mat but everyone on the same beat. Because every core is doing the identical step at the same moment, the hardware needs only one instruction decoder feeding a whole row of arithmetic units, which is exactly what frees up the transistor budget to have so many of them. The cost is rigidity: the cores move in lockstep, all on the same beat.

That single design decision is the source of both the GPU's power and its limits. When your work really is the same operation over a huge array of data, SIMD is a superpower. When your work needs each element to do something different, marching everyone in lockstep stops helping.

Many simple cores beat a few clever ones

same silicon budget, opposite layout core core core core CPU — a few big cores most silicon on caches & control GPU — thousands of tiny cores most silicon on arithmetic units
A CPU spends its silicon on a few large cores; a GPU spends the same silicon on thousands of small ones.

The core counts are not close. A modern consumer CPU has a few to a few dozen cores — roughly 4 to 24 on mainstream desktops and laptops, and up to around 64 on high-end parts, as of the mid-2020s. A modern GPU has thousands, from a couple thousand on modest cards to over ten thousand on high-end ones; a high-end board like the GeForce RTX 4090 lists 16384 of its cores.

The GPU can afford that many only because each one is stripped down. A CPU core carries deep caches and elaborate control logic so a single thread never stalls; a GPU dedicates most of its transistors to arithmetic instead, spending almost nothing per core on cleverness. Simple and numerous beats complex and few — but only for the right work.

That qualifier is the whole game. Thousands of cores are worthless if the job cannot be cut into thousands of independent identical pieces; then most of them sit idle while a handful do the work, and the CPU's few fast cores would have finished sooner. The GPU's advantage is real only when the work fills its width.

What GPUs are good for

The work a GPU loves has a name: embarrassingly parallel, meaning it takes little or no effort to split into a pile of independent tasks that need not talk to each other. When the same operation applies to millions of separate data elements, a GPU is in its element.

Graphics is the original case and the reason GPUs exist: every pixel on the screen can be shaded by the same little program, all at once. The same shape turns up far from graphics. Multiplying large matrices is thousands of identical independent multiply-add sums, which is why linear algebra runs so well on a GPU. Training neural networks — a later concept — is mostly that same matrix math, which is why the whole field runs on GPUs.

The common thread is uniform work over independent data: same operation, different element, no waiting on the neighbors. When a problem has that shape, handing it to thousands of cores turns hours into seconds.

What they're bad at, and how they team up

the CPU runs the program and offloads the parallel chunks CPU sequential & control GPU wide parallel work PCIe bus parallel batch results back sequential work stays here
The CPU runs the program and hands wide parallel batches to the GPU over the bus, reading results back, while sequential work stays on the CPU.

Give a GPU the wrong shape of work and its thousands of cores turn into a liability. A task that must happen in strict order cannot be spread across lanes at all — the GPU's width goes to waste while the CPU's fast single core would have finished it sooner. Heavy branching is nearly as bad: when the cores run in lockstep and one needs to take a different path than its neighbors, the hardware has to walk both paths in turn, so divergent code serializes the very parallelism it was meant to exploit.

So the two chips are not rivals; they are a team with a clear division of labor. The CPU runs the program, handles the branchy, sequential, decision-heavy work, and when it hits a big uniform parallel chunk it offloads that chunk to the GPU, then collects the results when the GPU is done. On a machine with a separate graphics card, that hand-off travels across a bus called PCIe.

The rule that falls out is worth keeping: reach for the GPU when the work is wide and uniform, stay on the CPU when it is twisty and sequential, and expect most real programs to need both.

Reading a GPU spec

Once the shape makes sense, a GPU spec sheet reads like plain language. The core count — thousands, versus the handful on a CPU — tells you how wide the chip is, how many identical pieces of work it can run at once. The memory figure tells you how big a batch of data can sit on the card at a time, close to those cores. Both numbers describe the same bet: enormous parallel width, fed by lots of nearby data.

This is why a self-driving RC car leans on a GPU. Reading a camera frame and deciding what is road, obstacle, or sign is a neural network doing the same arithmetic across every patch of the image — wide, uniform, and constant. That is exactly the work a handful of CPU cores would choke on and thousands of GPU cores devour, frame after frame, fast enough to steer.

The takeaway outlasts any one project: match the chip to the shape of the work. When you see a job that is the same operation over a mountain of independent data, you now know why the answer is a GPU — and why, when the job is one twisty sequence, it is not.