Training versus inference — study guide

The concept's fragments, read in order.

Learn once, use many times

training dataset model adjust parameters trained model fed in produces inference new input trained model prediction goes in produces
A model's two phases: in training a dataset flows into a model whose parameters are adjusted, producing a trained model; in inference a single new input flows into that trained model, producing a prediction.

A model does not live one kind of life. It has two distinct phases, and they feel so different that it helps to name them apart from the start. In training, the model learns its parameters from data: it works through examples and adjusts internal numbers until they capture the pattern. In inference, it uses those settled parameters to turn a new input into an output — a prediction, a label, an answer.

The split is the reason a model can be slow and costly to build yet quick and cheap to use. Training is the effortful part, done ahead of time and not on the clock of any one request like studying for an exam ahead of time — slow and effortful — and then sitting the exam quickly using what was already learned. Inference is the part that actually runs when someone needs an answer, over and over, long after the learning is done.

Everything in this concept follows from that one asymmetry: what the expensive learning involves, why serving predictions is comparatively light, where each phase runs, what hands off between them, and why a trained model does not stay correct forever. A model that fits parameters to examples is only half the story; the other half is putting those parameters to work.

The expensive learning

Training is where the model learns. It processes large amounts of data and repeatedly adjusts its parameters until those numbers capture the relationship between inputs and the answers it is meant to produce. The model does not receive its parameters ready-made; it arrives at them by working through the data, and getting them right is the whole job of this phase.

This is the compute-heavy part. Fitting parameters across a large dataset means an enormous amount of arithmetic, run again and again as the numbers are nudged toward better values. It is demanding enough that training is done occasionally rather than continuously — you pay for the learning in a concentrated burst, produce a model, and then leave it alone rather than relearning on every request.

That occasional rhythm is the point. Training happens offline, off to the side of whatever the model will eventually serve, so the heavy work never sits in the path of a user waiting for an answer.

Putting the model to work

Inference is the model earning its keep. It takes the already-trained model and runs it on a new input to produce an output — the prediction the whole exercise was for. Nothing is learned here; the parameters are fixed, and the work is simply applying them to an input the model has never seen before.

Because there is no learning to do, a single inference is fast. One input goes in, the fixed parameters do their arithmetic, and an answer comes out — quick enough to happen while a user waits. This is the phase that runs in production: it is what serves people, handles live requests, and does so continuously, not in the occasional bursts that training runs in.

So the model spends almost all of its working life in inference. It was trained once, but it answers again and again, and that steady stream of answers is the reason it exists.

One big bill, then many small ones

one big bill, then many small ones training one upfront cost inference a small cost per request running total grows with usage training's height, for comparison
Cost asymmetry: one large upfront training-cost block on the left, and on the right a long row of small per-request inference units whose running total climbs to rival and then exceed the training block.

The two phases do not just feel different — they cost differently, and in opposite shapes. Training is a large upfront cost. It ties up many accelerators — the GPU chips built for the heavy parallel arithmetic that learning demands — and it holds them for a long time, so the bill lands mostly at once, paid before the model has served anyone like setting up a printing press, a big one-time effort, after which each printed copy is cheap but adds up across a long run.

Inference costs the other way. Each request is comparatively cheap, but you pay it again on every single call, so the total is recurring and scales with usage. A model that answers a handful of times a day costs little; one that answers constantly runs its meter without stopping. The direction of the asymmetry is the thing worth internalizing: over a model's working life, the recurring inference bill can grow to rival or even exceed the one-time cost of training it.

That is why the phases have to be budgeted as two separate things. The upfront learning is a capital-sized decision made once; the ongoing serving is an operating cost that follows demand, and letting one stand in for the other is how the arithmetic goes wrong.

A cluster to train, less to serve

The two phases run on different hardware because they ask for different things. Training typically runs on a cluster of accelerators — many GPU chips wired together to share one large learning job — because fitting parameters across a big dataset needs that much parallel arithmetic held for that long. It is the kind of machinery you rent or build once and keep busy while the model is being made.

Inference asks for far less. Applying a fixed model to one input at a time can run on a single GPU, on a small edge device out where the data is, and for lighter models even on an ordinary CPU. What matters in this phase is latency: when someone is waiting on a real-time answer, the time each request takes is the thing to keep small, and small delays add up quickly across many calls.

So the shape of the deployment mirrors the shape of the cost. The heavy cluster does its work off to the side and occasionally; the lean serving side sits close to the request and runs whenever one arrives.

Weights are the handoff

Something concrete has to pass from the learning phase to the serving phase, and that something is the model's parameters — the weights. Training's entire product is this set of learned numbers: when the learning is done, the weights hold everything the model figured out from the data. They are the artifact, a data structure you can save to a file and carry around.

Inference is what loads those weights and runs them. It does not relearn anything; it reads the settled parameters, applies them to a new input, and produces an output. Because the knowledge lives in a saved artifact, the phase that made it and the phase that uses it can be cleanly separated — different time, different machine, different team even. Train here today, load the weights and serve them there next week.

That handoff is what makes the whole two-phase arrangement work. The expensive learning gets frozen into a portable object, and from then on using the model means loading that object rather than repeating the effort that produced it.

Models go stale

a model is never quite done deployed model drifts out of date retrain on fresh data over time triggers redeploy the updated model
The retraining loop: a deployed model drifts out of date over time, which triggers retraining on fresh data, and the updated model is redeployed back into service.

A trained model captures the world as its data described it, and the world does not hold still. As the inputs a model sees in production drift away from the data it learned on, its predictions drift out of date with them, and its accuracy quietly degrades like a map going out of date as roads change, so it must be re-surveyed now and then to stay accurate. Assuming a model trained on yesterday's data stays correct forever is a mistake with a name — it is one of the classic ways deployed models fail.

The fix is to treat a deployed model as something that ages rather than something finished. Teams watch its accuracy, and when the drift grows large enough they retrain it on fresh data — running the expensive learning phase again to fit new parameters — and then redeploy the updated model in place of the old one. The two phases turn out not to be a straight line from training to inference but a loop: serve, watch it go stale, retrain, serve again.

So a model is never quite done. Keeping it accurate means paying the training cost more than once, on a schedule the changing world sets rather than one you choose freely.

Budgeting the two bills

The moment you build anything on a model, the two phases stop being an abstraction and become two separate lines in the plan. A self-driving RC car learns to steer from data offline, once, and then runs inference in real time as it drives — the learning and the driving are different jobs with different budgets. The same split shows up whenever a model sits inside a larger system that has to keep responding to live requests.

Holding the phases apart is what keeps the reasoning straight. The training cost is a big number you commit to up front, and possibly again each time drift forces a retrain. The inference cost is a smaller number you pay on every prediction, so it grows with how much the thing is actually used. Confuse the two and you will either balk at a one-time bill that pays off over years or be blindsided by a per-request cost that quietly outgrows it.

So the load-bearing habit is simple: when a model is involved, ask what it costs to teach it and, separately, what it costs to run it — and budget both, because you will pay both.