Read this as a study guide instead

Loss and optimization

A score for being wrong, and turning the knobs downhill.

Loss curve
lossparameter

The same ideas, as prose

These are the exact fragments the model serves — also available as an ordered study guide.

Learning is just getting less wrong

A model fits its parameters to examples, but that raises an obvious question: fits them how? Something has to decide whether a given set of parameters is good or bad, and something has to change them toward good. Those two jobs are the whole of how a model learns, and they split cleanly in half.

The first half is a loss function, a single number that scores how wrong the model's predictions are against the known answers — the smaller the number, the better the fit. The second half is an optimizer, the procedure that adjusts the parameters to push that number down. One measures, the other moves.

Everything in this concept is those two halves worked out in detail: what the score actually measures, how the optimizer knows which way to nudge each parameter, how big its steps should be, and what it means to run the process until the number stops falling. Learning, stripped to its mechanism, is just getting less wrong on purpose.

A number for how wrong

Before a model can improve, its wrongness has to become measurable. A loss function does exactly that: it takes the gap between what the model predicted and the answers you already know, and it collapses that gap into one number like a single score that tallies how far off you are, which you keep trying to drive toward zero. Lower is better, and a loss of zero would mean the predictions matched the known answers exactly.

The point of squeezing everything down to a single value is that it gives learning a concrete target. A cloud of individual errors across many examples is not something you can steer toward; one number is. The whole of training becomes a single, unambiguous instruction: make this number smaller.

That number is also the only channel through which the data speaks to the parameters. The model never sees "right" or "wrong" directly — it sees a loss, and its entire sense of how it is doing comes from whether that score is high or low.

Find the parameters that minimize it

loss over the space of parameters loss parameter lowest loss: the target
A U-shaped loss curve plotted against a parameter axis, with its single lowest point marked as the target the optimizer is trying to reach.

Once wrongness is a number, learning turns into a search. Optimization is the hunt for the parameter values that make the loss as small as possible — the setting of the model's internal numbers at which it is least wrong about the examples it has seen.

It helps to picture the loss as a landscape. Spread every possible choice of parameters out as a surface, and let the height at each point be the loss you would get with those parameters. High ground is a badly wrong model; the low points are the good ones. Optimization is the problem of finding a valley in that surface without being able to see the whole map at once.

That framing is what the rest of the mechanism rests on. The optimizer stands at some point on the landscape, decided by the parameters it currently holds, and its only task is to get lower.

Step downhill, repeat

feel the slope, step the opposite way loss parameter slope (gradient) current step downhill next: closer to minimum
A point on a loss curve with its tangent slope drawn, and a step arrow moving the point downhill in the direction opposite the slope, landing closer to the minimum.

The optimizer cannot see the whole landscape, but it can feel the ground right under it. At the current parameters it computes the gradient — the slope of the loss, which way is uphill and how steeply. That local slope is all the information it needs to take one good step.

Gradient descent is the rule that uses it: move the parameters a step in the direction opposite the gradient, which is downhill, toward lower loss like walking downhill in thick fog, feeling the slope underfoot and stepping the way the ground falls, again and again. Then measure the slope again at the new spot and step again. Each step lowers the loss a little, and repeating the measure-and-step loop walks the parameters steadily toward a valley. The process runs until the loss stops falling meaningfully — the point called convergence, where the ground has gone flat and further steps buy almost nothing.

The direction is the part worth pinning down: descent means stepping against the slope, not along it. A step that followed the gradient would climb, raising the loss, which is precisely the wrong way.

How big a step to take

the learning rate sets how far each step moves loss too small: tiny steps barely moves toward the bottom just right: settles reaches the minimum too large: overshoots climbs away and diverges
Three descents on the same loss curve: a too-small learning rate takes tiny steps that barely move, a good rate settles at the minimum in a few steps, and a too-large rate overshoots back and forth and climbs away.

Knowing which way is downhill leaves one thing undecided: how far to move each step. That distance is the learning rate, and it is the single knob that most shapes whether training works like the length of your stride: tiny steps take forever to cross the ground, while huge leaps fly right past where you meant to land.

Both extremes fail, in opposite ways. Set the learning rate too small and each step barely moves the parameters; the loss does fall, but so slowly that reaching a valley takes an unreasonable number of steps. Set it too large and the steps overshoot the low point, landing further up the far side than they started; do that repeatedly and the loss climbs instead of falling, a runaway called divergence. A workable rate lives between the crawl and the overshoot, large enough to make real progress yet small enough to settle rather than bounce.

There is no single right value that works everywhere, which is why choosing the learning rate is one of the central practical problems of training a model.

Many steps toward a bottom

One step lowers the loss a little; training is what happens when you repeat that step over and over. Iteration after iteration, the parameters inch downhill and the loss decreases, until the steps stop making a difference and the process has settled at a minimum — the bottom of the valley it happened to be walking down.

The catch is that a landscape can have more than one valley. A loss surface that is smoothly bowl-shaped has a single lowest point, and descent is guaranteed to head for it. Many real loss surfaces are not that kind: they are bumpy, with several dips of different depths. A minimum that is the bottom of its local dip but not the deepest dip anywhere is a local minimum, distinct from the global minimum that is the true lowest. Gradient descent only ever feels the slope right under it, so it can walk into a shallow dip, find flat ground, and stop — having minimized the loss locally without reaching the lowest point on the map.

Reaching a low training loss is the goal of this machinery, though a lower training loss is not on its own a guarantee of a better model in the world. What the iterations reliably deliver is a bottom; which bottom, and how good it is, is the harder question.

A little data at a time

Computing the exact gradient means running every training example through the model before taking a single step. With a large dataset that is enormously expensive, and it buys one precise step at the cost of touching all the data every time.

Stochastic gradient descent, and its common form using minibatches, makes a cheaper trade. Instead of the whole dataset, it draws a small random batch of examples and estimates the gradient from just those. The estimate is rougher than the full-dataset gradient, but it is far cheaper to compute, so the optimizer can take many quick steps in the time one exact step would have taken. Across many steps the rough estimates average out and still carry the parameters downhill.

The roughness even turns out to be useful. Because each batch is a different random sample, the estimated slope jitters from step to step rather than pointing the same exact way, and that noise can jostle the parameters out of a shallow local dip they would otherwise have settled into. The approximation that was supposed to be a compromise sometimes finds a better bottom than the exact gradient would.

The engine under every trained model

Loss and gradient descent are not one technique among many; they are the engine that turns data into a working model. Nearly every trained system you encounter got that way by the same loop — score the wrongness, feel the slope, step downhill, repeat — running on a scale that is hard to picture. The specific architecture on top varies enormously; the machinery underneath rarely does.

It is also the point where a vague ambition becomes a concrete process. "Get better at driving" is not something a machine can act on, but "make this loss number smaller" is. A self-driving RC car learns exactly this way: its errors on recorded driving become a loss, and gradient descent nudges its parameters, step after step, until the number is low and the car keeps to the track. The wanting-to-be-good is turned into a slope to walk down.

That is why this concept is load-bearing for everything built on learned models. Understanding how a model gets better means understanding a loss going down and an optimizer pushing it there — and once you see that loop, you can see it running underneath the whole field.