Fine-tuning versus prompting — study guide
The concept's fragments, read in order.
Change the input, or change the model
You have a capable pretrained model and a task it was not built specifically for. There are two ways to bend it toward what you want, and they act on two different things. You can change the input — write the request carefully so the model behaves the way you need — or you can change the model itself, continuing its training on your own examples until its weights carry the task.
The first lever is prompting: it works at inference, when the model is just running, and it leaves the weights exactly as they were. The second is fine-tuning: it is more training, and training is the part that changes weights. The distinction is the whole of this concept like giving a skilled worker clear instructions for each job versus sending them on a course that changes how they work by default — one steers a fixed model from the outside, the other rebuilds a little of the model on the inside.
Naming which lever you are pulling settles almost everything that follows: what it costs, how fast you can try it, whether you can undo it, and when the heavier option is worth reaching for. Most of the time you will pull one of these two, and often, in the end, both.
Steering at inference time
Prompting shapes what a model does purely through the input you hand it. You describe the task, set the tone, spell out the format you want, and the model responds accordingly — all at inference, the moment it is running to answer a request. Nothing about the model is rewritten in the process.
That is the defining fact: the weights never change. The pretrained model that answered your last request is the same model that answers this one, steered only by different words. Because there is no training step, the effect is immediate — you edit the request and the behavior shifts on the very next run — and it is fully reversible, since going back to the old behavior means going back to the old wording.
This makes prompting the cheap, fast place to experiment. You can try one phrasing, see it fall short, and try another in seconds, without a dataset, a training job, or a new model to manage. The limit is that you are always working from outside a fixed model: you can direct the capabilities it already has, but you are not adding new ones to its weights.
Adapting the weights
Fine-tuning takes a pretrained model and keeps training it, this time on examples you provide — inputs paired with the outputs you want back. Training is the part that adjusts weights, so this is not steering a fixed model from the outside; it is changing the model itself so the task settles into its parameters.
What you get out is a new, specialized model rather than a new prompt. You show it enough of the kind of work you care about, the training run nudges the weights toward that work, and the result is a model that leans your way by default. The task lives in the weights now, not in the instructions you attach to each request.
That is the trade at the heart of it. Prompting leaves a general model general and redirects it request by request; fine-tuning spends a training run to bake your task in, producing a model that no longer needs to be told the same thing every time. The gain is a model shaped to you; the cost is that you had to build one.
What each one costs
The two levers do not just cost different amounts; they charge you at different times. Prompting has no training cost at all and works immediately — there is nothing to build before you start. What it does cost, it costs on every request: the instructions and any examples you put in the prompt ride along with each call, so you pay for them again each time the model runs.
Fine-tuning inverts that. Its cost is paid up front, before the model serves a single request. You have to curate training data — assemble the inputs and the outputs you want the model to learn — and then pay for a training run to actually adjust the weights, which takes time. After that you have a tuned model to host and keep serving, an ongoing cost of its own.
So the shape of the bill differs as much as the size. Prompting is a small charge repeated forever; fine-tuning is a large charge once, plus keeping the result running. Which is cheaper depends less on the price of either and more on how much you will use the thing — a small charge paid enough times eventually overtakes a single large one.
Start by asking
For most tasks, the right first move is to prompt. It is the fastest way to find out whether the model can already do what you need, and often it can — a careful request gets you far enough that nothing heavier is called for like rewording a question to get a better answer before assuming the person needs retraining. Since prompting changes nothing about the model, trying is nearly free, so there is little reason not to try it first.
Prompting also has more range than a single instruction suggests. You can put a few examples of the output you want right into the prompt — giving the model a handful of worked cases to imitate, a practice called few-shot — and steer it further still without any training. That keeps the whole loop tight: phrase, run, look at the result, rephrase.
The discipline is to exhaust this cheap iteration before spending on anything expensive. You learn what the model can and cannot be talked into, you refine the request against real outputs, and only the gaps that survive all that phrasing are honest evidence that you need more than a prompt.
When asking is not enough
Fine-tuning earns its up-front cost when prompting has been tried and still cannot get you there. A few needs tend to push you across that line like sending someone to a specialist course when no amount of in-the-moment instruction gets the consistency you need. The first is consistency: when you need the model to hold a specific output format or style reliably, request after request, rather than mostly honoring an instruction. The second is depth: adapting the model to a specialized domain whose patterns are hard to convey in the space of a prompt.
The third reason is about scale rather than quality. A fine-tuned model has the task baked into its weights, so it needs far less spelled out in each request — the prompts get shorter. At high volume, those shorter prompts mean lower cost per request and can mean lower latency too, which is exactly the case where paying once up front beats paying a little on every one of many calls.
The common thread is that fine-tuning is the answer to a demonstrated limit, not a default. You reach for it when a real gap in consistency, domain fit, or per-request economics has survived honest prompting — not before you know one exists.
Not either/or
Prompting and fine-tuning are not rivals you must choose between once and for all. A fine-tuned model is still a model you drive with a prompt, so the natural setup is to do both: fine-tune to move the heavy, consistent behavior into the weights, then prompt the tuned model to handle the specifics of each request. The lever you pulled first does not lock out the other.
There is also middle ground between full fine-tuning and pure prompting. Lighter-weight, parameter-efficient methods — LoRA and other adapters, grouped under the name PEFT — adapt a model by training only a small piece of it instead of retraining everything. They are worth knowing exist as points on the spectrum between the two ends; the mechanics are a separate matter.
The takeaway is that the two levers describe a range, not a fork. Real systems mix them, sliding from a bare prompt toward heavier adaptation only as far as the task demands.
Prototype cheap, specialize when forced
The reason this distinction is load-bearing is that it decides the order in which you spend. Prompting costs almost nothing to try and tells you quickly whether a model can do your task at all, so it is the right way to prototype. Fine-tuning costs data, time, and a training run before it returns anything, so it is the wrong way to start.
Reaching for the heavyweight path too early is a real and common mistake. You can pour effort into curating examples and running a training job to solve a problem that a better prompt would have solved for free — spending the expensive resource before you had evidence you needed it. The cheap experiment exists precisely to keep you from that.
So the working rule is plain: prototype by prompting, and specialize by fine-tuning only when a demonstrated limit forces you to. Knowing which lever you are pulling, and pulling the cheap one first, is most of what it takes to spend your data and time where they actually buy something.