← BLOG
October 06, 2026ai-agentsenterprise3 min readITENHR

Which prompt made that decision?

Most teams cannot say which prompt, model version and retrieval index produced a given AI decision. Treat all three as versioned release artifacts and record them with every decision.

Three versioned cards labelled prompt, model and index, feeding a decision record that stores all three

Pick any decision your AI feature took in production three weeks ago and ask a simple question: which prompt produced it? In most teams the honest answer is "probably the current one, unless someone changed it". That is not an answer you can give to an incident review, to a customer, or to an auditor.

Three things change behaviour, and none of them is in the release

An AI feature's behaviour depends on more than its code. The prompt decides how the model is instructed. The model decides how those instructions are interpreted. The retrieval index decides what context the model sees. Change any of the three and the same input can produce a different decision.

In many systems none of the three goes through a release. The prompt lives in a database row or an admin panel, edited directly in production. The model is referenced by an alias such as "latest", and the provider moves that alias to a new version on its own schedule. The index is rebuilt by a nightly job. Nobody breaks anything, yet behaviour drifts, and when someone asks why, there is nothing to compare against.

Treat them as release artifacts

The fix is not a new tool. It is the release discipline we already apply to code, extended to the parts that change behaviour most.

The prompt is code. It lives in the repository, goes through review, has a version and ships with a release. The model is a dependency: pin the exact version identifier the provider exposes, never the alias, and upgrade it deliberately, like any other dependency. The retrieval index gets a version too, tied to the documents and the embedding model that built it.

Every change to any of the three goes through the same gate: run the eval set against the new combination, compare it with the one in production, then release. Rollback means redeploying the previous combination, not retyping a prompt from memory or hoping the provider still serves the old model under the old name.

Record the triple with every decision

The last step is what makes the rest useful. Every decision the system takes records the triple it ran on: prompt version, model identifier, index version. Put it next to the input and the output in whatever log you keep.

With that, an incident review starts from facts: this decision, this prompt, this model, these documents. You can reproduce it, compare it with current behaviour and say whether a change caused it. Without it, every answer about past behaviour is a guess, and so is every answer to someone who has the right to ask.

None of this is exotic. It is versioning, pinning and release gates. We just forgot to apply them to the part of the system that changes behaviour the most.

Where does the prompt you run in production live today, and could you say which version took last month's decisions?