You probably don't need fine-tuning
Fine-tuning is often the first fix proposed when an AI feature underperforms. Most of the time the model lacks context, not capability — and without an eval set you cannot even tell whether tuning helped.

When an AI feature underperforms, fine-tuning is often the first remedy on the table. It sounds like serious engineering: a dataset, a training run, a model that is "ours". In practice it is frequently the most expensive way to avoid looking at why the feature underperforms in the first place.
The model rarely lacks capability
In the failures I see, the base model is usually capable of the task. What it lacks is context. The retrieval step returns the wrong documents, or the right ones in the wrong order. The instructions have accumulated over months and now contradict each other. The input arrives half-parsed: a PDF read badly, a table flattened into a paragraph, a field missing. The model is asked to answer a question it cannot see properly.
None of that is fixed by changing the weights. A tuned model fed the same broken context produces the same broken answers, only with more confidence and a higher bill.
Fine-tuning freezes a snapshot
A fine-tuned model learns the process as it was when the dataset was built. The process keeps moving: the catalogue changes, a regulation changes, the customer adds an exception. The tuned model does not follow. Every meaningful change means a new dataset, a new training run and a new validation.
It also ties you to a specific base model version. Model providers release new versions regularly; a tuning built on the previous one has to be redone or abandoned, and prompt-level improvements that transfer for free do not come with it.
No eval set, no fine-tuning
The decisive argument is simpler. Without an eval set — a fixed collection of real cases with expected outcomes, run on every change — you cannot tell whether the tuning helped. You will have an impression, a few good demos, and a cost.
If you have no eval set, you are not ready to fine-tune. If you have one, run it against better prompts and better retrieval first. Often that closes most of the gap, and you will know exactly how much remains.
When it does make sense
Fine-tuning earns its place on narrow, stable, high-volume tasks: a fixed output format, a classification that does not change every month, a job you want to move to a smaller and cheaper model once the large one has shown what good looks like. In those cases the evals tell you so, with numbers you measured yourself.
The order matters: prompt, retrieval, eval set, then — maybe — the weights.
What was the last problem your team tried to solve with fine-tuning, and what did the evals say?