Reading a live system — study guide

The concept's fragments, read in order.

What a machine is doing right now

A running machine is always spending something. Every instant it is dividing a fixed amount of processor time, a fixed amount of memory, and a fixed rate of disk traffic among whatever it is currently doing. None of the three is infinite, and reading how much of each is in use is how you tell what the machine is doing and whether it is in trouble.

Three resources carry almost every performance story. The CPU does the computing; memory (RAM) holds what the running programs need close at hand; disk I/O is the traffic to and from storage. A machine that feels slow is nearly always short on one of the three, and the same small set of read-only tools reports all three on any operating system.

Reading a resource is not the same as changing it. Everything in this concept is diagnosis — looking at live numbers a machine already keeps about itself. Deciding what to do about a number, killing a process or freeing space, is a separate and deliberate act that comes only after you know which resource is the problem.

CPU usage and load

CPU utilization and load average sound like the same measurement and are not. Utilization is the fraction of time the processor spends busy rather than idle, and on a multi-core machine it is reported per core — a summary line typically splits that time into user, system, idle, and I/O-wait shares, and each process gets its own percentage of the total CPU time it used.

Load average is the other number: the three figures a command like uptime prints at the end of its line, the average number of tasks in the run queue over the last 1, 5, and 15 minutes. It is not a percentage, and it is not scaled by the number of cores, so you read it against the core count yourself. A load of 1 on a four-core machine means the box sat idle roughly three-quarters of the time, while a load above the core count means work is queueing for cores that are already busy.

One nuance saves a lot of misdiagnosis. On Linux the load average counts not only tasks running or waiting for the CPU but also tasks stuck waiting on disk (the uninterruptible state), so a high load average does not always mean the CPU is the problem — heavy disk I/O alone can drive it up while the processor sits mostly idle. When load is high and utilization is not, look at the disk.

Used, free, and the cache in between

one RAM bar: used, cache, and free used by programs reclaimable cache free available = free + reclaimable cache (the kernel's estimate of room for new programs) a low free is normal; available is what tells you if you are short
A single RAM bar splits into three regions - memory used by programs, reclaimable file cache, and free memory - and available memory is the free region plus the reclaimable cache, so a low free number is not a low available number.

Memory reporting has a trap in it: the most alarming-looking number is usually the least important one. A tool like free shows memory as used, free, and available, and beginners fixate on a tiny free figure and conclude the machine is out of RAM. It almost never is.

The reason is that the operating system refuses to let RAM sit empty. It fills otherwise-idle memory with a reclaimable cache of files recently read from disk (often shown as buff/cache), because cached RAM answers a repeat read instantly while empty RAM does nothing at all — like a chef keeping ingredients out on the counter within reach, then sweeping them aside the instant the space is needed for something else - an empty counter was never the goal. That cache is not really spoken for: the moment a program needs the space, the kernel takes it back.

So the number that actually tells you whether you are short on memory is available, not free. Available is the kernel's estimate of how much memory a new program could take without forcing the system to swap — free memory plus the cache it can reclaim. A machine can show almost no free and still have most of its memory available, which is exactly the healthy state; it is when available runs toward zero that you are genuinely out of room.

When memory runs out — swap

when RAM fills, a cold page spills to disk RAM (full) evict a cold page fetch it back - slow swap area (on disk) far down the memory hierarchy
RAM is full, so the operating system evicts a cold page down to a swap area on disk; because disk sits far down the memory hierarchy, fetching that page back into RAM is slow.

When memory genuinely fills — when available, not just free, runs low — the operating system needs room it does not have, so it makes some. It takes pages of memory that have not been touched in a while (cold, inactive pages) and writes them out to a reserved area on disk called swap, freeing that RAM for whatever needs it now.

The catch is where swap lives. Disk sits far down the memory hierarchy, and the gap between RAM speed and disk speed is enormous, so reading a page back from swap is dramatically slower than reading it from memory. A little swapping is unremarkable. But when a machine is so short on RAM that it keeps evicting pages it immediately needs again, it falls into thrashing: it spends its time shuttling pages between RAM and disk instead of doing work — like a desk too small for your papers, so you keep carrying stacks to a filing cabinet down the hall and back - past a point, all the walking leaves no time for the work. The tell is distinctive and almost paradoxical: the CPU is barely busy while paging activity is enormous, because the processor is mostly waiting on the disk.

Thrashing is why a machine short on memory can feel more broken than one short on CPU. A busy CPU is at least making progress; a thrashing machine is running hard and getting almost nothing done.

Disk I/O and being blocked on it

Input/output — I/O — is the traffic between a program and storage: reading files in, writing results out. It matters to resource-reading because of what a process does while it waits. A process waiting on the disk is blocked: it is not using the CPU at all, just parked until the data arrives. This is the blocked state every process moves through, and it means a machine can look busy while its processor is idle, stalled behind slow storage rather than computing.

That is also why disk pressure can masquerade as CPU pressure. Because tasks blocked on disk still count toward the load average on Linux, a machine can post a high load while the CPU sits mostly idle — the queue is full of processes waiting on I/O, not fighting over cores.

Two different disk problems hide under the word full, and telling them apart matters. One is space: the filesystem has no free bytes left, so writes fail outright, and a df report shows the disk at capacity. The other is throughput: the disk has plenty of free space but cannot move data fast enough to keep up, so everything that touches it slows down. A disk can be starved for either one while healthy on the other, and the fix for each is different.

Reading a live monitor

A live monitor like top on Unix or Task Manager on Windows is one idea repeated on a timer: it takes a snapshot of the whole machine, displays it, and a moment later takes another. What you watch is not a live wire into the hardware but a sequence of still frames refreshed fast enough to look continuous.

Every such monitor has the same two-part shape. At the top is a system summary — overall CPU utilization, memory used and available, and the load average — describing the state of the machine as a whole. Below it is a per-process list, one row per running program, each with its own share of the CPU and its own slice of memory. The summary tells you whether there is a problem; the list tells you which process owns it.

That pairing is the entire skill. A summary showing memory nearly exhausted, read together with the process sitting at the top of the memory column, turns the vague question of whether the machine is in trouble into the specific one of which program to look at next.

Is it CPU, memory, or I/O?

one slow machine, three answers - the reading tells you which CPU-bound memory-bound I/O-bound CPU pinned near 100% load above the core count work queued for busy cores available memory near zero swap in active use thrashing: low CPU, heavy paging processes blocked on disk high load, low CPU long disk waits more CPU, not more RAM more RAM, not more CPU faster or less disk I/O
A slow machine is CPU-bound, memory-bound, or I/O-bound, and the live reading tells them apart: the CPU pinned near 100 percent, available memory near zero with swap active, or processes blocked on disk with high load but low CPU.

When a machine is slow or stuck, the useful question is not how slow but which resource ran out — and there are only three answers worth checking first. The machine is CPU-bound, its processor pinned near 100% with work queued behind it; or memory-bound, out of usable RAM and swapping to disk; or I/O-bound, its processes blocked waiting on storage while the CPU idles. Each has a distinct signature in the numbers, so the live readings sort one from another.

The reason this is worth drilling is that the failures you will actually meet almost always land in one of the three. A game server that lags under a full lobby is a fine example: it is nearly always short on CPU to simulate the world, short on memory and paging, or stalled on disk — and which one it is decides the fix entirely. Adding RAM does nothing for a CPU-bound server, and a faster CPU does nothing for one that is thrashing.

So the payoff of reading a live system is triage. You do not guess and you do not start changing things blindly; you look at CPU utilization and load, at available memory and swap, at disk wait and free space, and you let the numbers tell you which of the three problems you have before you touch anything.