Governability: who changed the prompt, and can you roll it back?
An AI feature is governable when you can say which prompt, model and settings produced any output, who approved that version, and how to go back to the previous one. How to build that in from the first release.
In most AI products, the behaviour that matters most lives in places that are easy to change and hard to track: a prompt in a settings table, a model name in an environment variable, a threshold someone adjusted in an admin screen. One small edit can change what thousands of users see, and a week later nobody can say what changed, when or why.
If you lead the team: what to ask
- For any AI output, can we tell which prompt, model and settings produced it?
- Who has to approve a change to how the AI behaves, and is there a record of it?
- If a change goes wrong, how fast can we go back to the last version that worked, and have we tried it?
Governability is the ability to answer three questions at any time: what produced this output, who changed it, and how do we go back? It is not paperwork added at the end. It is a set of engineering choices that cost little at the start and a lot to retrofit.
Version everything that shapes behaviour
Treat prompts, model choices and AI configuration as versioned artefacts, the same way you treat code:
- Prompts live in the repository or in a versioned store, each with an ID, never as free text edited in place.
- Model and provider are pinned to a specific version, not “latest”, so a provider update does not silently change your product.
- Configuration such as temperature, tools allowed, thresholds and rule weights is versioned with the prompt it belongs to.
A release of an AI feature is then a known combination: this prompt version, this model, this configuration.
Provenance on every output
Every stored output should carry its origin: which prompt version, which model, which scorer or rule set, and when. When a manager asks why a score looks wrong, you can see exactly what produced it. When a version turns out to be flawed, you can find every output it produced and decide whether to re-run, flag or correct them. Without provenance, a bad version leaves results behind that you cannot separate from the good ones.
Review changes like code
A prompt change can alter behaviour more than a hundred lines of code, so it deserves the same discipline:
- a written reason for the change;
- a diff someone else reads before it ships;
- a run against a fixed set of test cases, including past mistakes, with results compared to the current version;
- a staged release, so the new version meets a small share of real traffic first.
Approvals for high-impact behaviour
Not every change needs a committee. Rewording a summary prompt is low risk. Changing how customers are scored, what the AI is allowed to do on its own or what it says about pricing, health or legal matters is not. Classify AI behaviours by impact and require an explicit approval from a named owner for the high-impact ones. Keep the list short and clear, so the rule is followed instead of worked around.
An audit trail people can read
Record who changed what, when, with which approval, and what the tests showed. Store it where it cannot be quietly edited. This is what lets you reconstruct a decision months later, and it is what compliance and security reviewers ask for first.
Rollback in minutes, rehearsed
If going back means a code change and a full deploy, rollback will be slow exactly when you need it fast. Make the active version a setting that can be switched back to the previous known-good combination in one step, and try it before you need it. Also decide what happens to outputs produced by the bad version: re-run, flag or leave with a note.
Documentation for compliance teams
Compliance, legal and security teams do not need the prompt text explained line by line. They need a short, current document per AI feature: what it does, what data it reads, what it can and cannot do, who owns it, how changes are reviewed and how it is monitored. Generate as much of it as you can from the versioned configuration, so it does not fall out of date.
One owner per AI behaviour
Every AI feature needs a named owner who decides on changes, watches its quality signals and answers for its behaviour. Shared ownership usually means nobody notices when it drifts.
Governability checklist
- Prompts, model versions and AI configuration are versioned together.
- Every output records prompt, model, scorer and time.
- Prompt changes are reviewed, tested against known cases and released in stages.
- High-impact behaviours need approval from a named owner.
- An audit trail records who changed what, when and why.
- Rollback is one step and has been rehearsed.
- Each AI feature has a short, current document and one owner.
If you are planning AI on top of your own data, our AI on your data service builds these controls in from day one.
Written from our engineers’ work on production systems. Want a second opinion on your project? Talk to an engineer.
See the work →