Observability — study guide

The concept's fragments, read in order.

Seeing inside a running system

A running process is a black box from the outside. It is doing work right now — serving requests, moving data, occasionally getting stuck — but nothing about that is visible unless the program says so. The old way to find out was to log into the machine and poke around by hand, hoping the problem was still happening when you got there. Observability is the alternative: inferring what a system is doing from the signals it emits, rather than guessing.

The word names a property of the system, not a tool you bolt on afterward. A system is observable to the degree that its outputs let you answer questions about its insides — including questions nobody thought to ask when it was built. That only works if the program was made to emit useful signals in the first place.

Those signals come in three standard kinds: logs, metrics, and traces. Each answers a different question, and the rest of this concept is what they are, what each one is good for, and what it costs to collect them.

The event trail

A log is a timestamped record of a discrete event: one thing that happened, and when. A request arrived, a file was opened, a payment failed — each becomes a line stamped with the moment it occurred. Read enough of them in order and you can reconstruct exactly what the system did, one step at a time like a flight recorder — a detailed, timestamped trail of events you read back afterward to reconstruct exactly what happened.

There are two ways to write that line, and the difference matters more than it looks. An unstructured log is free-form text — a sentence a human wrote for another human to read. A structured log carries the same information as typed key-value fields: a level, a timestamp, a request id, a duration, each in its own named slot. Structured logs are far easier to search and aggregate than free-form text, because a machine can filter and group them by field instead of guessing at the shape of a sentence.

The catch is that a busy system emits an enormous number of these records, and reading them all is not an option. Logs tell you what exactly happened once you know roughly where to look — which is why they work best alongside the other two signals, not alone.

Numbers over time

a metric sampled over time, and the threshold that trips an alert value time threshold alert fires value crosses the line
A metric plotted as points over time rises until it crosses a horizontal threshold line, and the crossing point is marked as the moment an alert fires.

A metric is a numeric measurement sampled over time: not a record of one event, but a number you watch move. Request rate, error rate, memory in use — each is read at intervals and stored as a series of values against the clock. Where a log is a sentence, a metric is a gauge you glance at to see the system's current state and how it is trending like the gauges on a car's dashboard — continuous numbers you glance at to see the current state and which way it is trending.

Metrics come in a few standard shapes. A counter is monotonically increasing — it only goes up, and resets to zero when the process restarts; total requests served is a counter. A gauge goes up and down, tracking a value that rises and falls, like the number of connections open right now. A histogram counts observations into buckets, so instead of a single number you get a distribution — how many requests finished quickly, how many slowly, and how many were painfully slow.

Metrics are cheap to store and quick to aggregate, which makes them the raw material for dashboards and alerts. Their weakness is the flip side of their strength: a metric tells you a number changed, but never which individual event caused it.

Following one request

one request as nested timed spans across services time api-gateway handle request auth-service auth check orders-service load order (slowest span) database db query total request duration the longest span is where the time went
One request drawn as a waterfall of nested timed spans across three services along a time axis; the load-order span is the longest and accounts for most of the request's total duration.

When one request crosses several services before it answers, no single log or metric shows the whole journey. A distributed trace does: it follows one request across services as a set of timed spans, where each span is one unit of work — a database query, a call to another service, a stretch of computation — stamped with when it started and how long it took. The spans nest into a tree, and together they show the request's end-to-end path like tracking a package — every hop along the way is stamped with a time, so you can see exactly where the journey slowed down.

Because every span is timed, a trace shows where the time actually went. A request that takes two seconds is rarely slow everywhere; it is usually one span — one downstream call, one query — eating most of the total while the rest wait on it. The trace makes that span obvious instead of leaving you to guess which service to blame.

That is the question traces answer and the others cannot: not whether something is slow, and not what a single service logged, but where in a chain of services the time was spent.

Which signal answers which question

three signals, three questions logs metrics traces timestamped events number over time nested spans what happened? is something wrong? where?
Three columns for the pillars of observability: logs shown as a list of timestamped events answering what happened, metrics shown as a time-series chart answering is something wrong, and traces shown as a span waterfall answering where.

Logs, metrics, and traces are often called the three pillars of observability, and the reason there are three rather than one is that each answers a different question. They are complementary, not competing: a system instrumented with only one of them can see only part of what it is doing.

Metrics tell you that something is wrong. An error rate climbs or a latency number spikes, and a cheap, always-on number is what catches it. Traces tell you where it is wrong — which service or which span in a multi-service request is the one burning the time. Logs tell you what exactly happened — the specific event, with its details, once the trace has pointed you at the right place and moment.

The workflow falls out of that division of labor: a metric raises the alarm, a trace narrows the search to one span, and a log explains the particulars. Skip a pillar and you feel the gap — a number with no story behind it, or a story you can never find in time.

From signals to attention

Collecting signals is worthless if nobody looks at them, and nobody can watch a number every second. Dashboards and alerts are the two ways signals reach a human: one for when you are already looking, one for when you are not.

A dashboard visualizes metrics — request rate, error rate, latency drawn as lines over time — so the state of a system is an at-a-glance picture instead of a query you have to run. It is what you open when you already suspect something, or just to keep an eye out. An alert is the opposite posture: it fires automatically when a value crosses a threshold or an SLO is at risk. An SLO, a service-level objective, is a target value or range for some service level — say, a ceiling on error rate or a floor on the fraction of fast requests — and an alert warns you before the system drifts past it.

The point of the alert is that it watches so you do not have to. Set the threshold too loose and you miss the problem; set it too tight and it cries wolf until you stop listening. Getting that line right is most of the craft.

Why you cannot measure everything

It is tempting to attach every detail you might one day want to a metric — the method, the status code, the endpoint, the customer. Doing so is how observability bills quietly balloon, because of how metrics are stored. Each unique combination of label values becomes its own stored time series, so labels multiply rather than add.

The multiplication is the trap. Imagine a single request counter labeled by method, status code, and endpoint: a handful of methods times a handful of status codes times a few dozen endpoints already stands for a large number of separate series, and this is only an illustrative example of the arithmetic. Now add a label with one value per customer, and that number is multiplied again by the customer count — a high-cardinality label like a user id or a full URL can turn one tidy counter into an unmanageable pile of series, each carrying its own storage and query cost.

Traces face the same pressure from a different direction, and the usual answer is to sample them: keep a representative subset rather than every trace, trading some completeness for a cost you can afford. So observability is never measure everything; it is a deliberate choice of what is worth the coverage.

Instrument before you need it

Observability is load-bearing at the exact moment you have no time for it. A game server buckling under a flash crowd, or an agent pipeline that has quietly started returning garbage, is debuggable only if it was already emitting the signals that explain it. You cannot go back and add a log line to an event that happened an hour ago.

That is the one habit worth taking from this concept: instrument before you need it, not after. Emit structured logs, count and time the things that matter, and trace the requests that cross services — while the system is healthy and you have the luxury of thinking clearly. The signals cost a little to collect all the time so that they are there the one time everything is on fire.

The three pillars are how you get that reach: metrics to notice, traces to locate, logs to explain. A system that emits them is one you can reason about under pressure; one that does not leaves you back at logging into a box and hoping.