Prompts are code: version them, test them, review them
A prompt decides what your AI system does, so change it the way you change code: in version control, reviewed, tested against fixed examples, and recorded on every output so you can explain it and roll it back.
In many AI products the most important logic isn’t in the code. It’s in a paragraph of English pasted into a settings screen, edited by whoever last had an idea, with no history and no tests. Change one sentence and every answer the product gives can change with it.
If you lead the team: what to ask
- Where do our prompts live, and can we see who changed them, when and why?
- Is a prompt change tested against fixed examples before it goes live?
- If a prompt change makes answers worse, can we go back in one step?
A prompt decides what your system does. Treat it like code: version it, review it, test it, and know exactly which version produced each output.
Keep prompts in version control
Store prompts as files in the same repository as the code that uses them, not in a database field or a vendor dashboard. Every change then has an author, a date, a diff and a reason. When an answer changes, you can see what changed in the prompt, and when.
Use templates with named variables
Build prompts from templates with explicit, named variables (the user’s question, the retrieved passages, the output format) instead of gluing strings together across the codebase. A template can be read in one place, rendered in tests with sample values, and checked: a missing variable is an error before release, not a strange answer in production.
Review prompt changes in pull requests
A prompt change goes through a pull request like any other change, reviewed by someone who understands both the product and the model. The reviewer asks the same questions as for code: what does this fix, what could it break, and where is the test? Small wording changes deserve this most, precisely because they look harmless.
Test against a fixed example set before release
Keep a fixed set of example inputs with expected outputs or acceptance criteria: typical cases, edge cases, and inputs that used to fail. Run the set on every prompt change and every model change, and compare with the last accepted run. Check automatically what you can: valid structure, required fields, content that must never appear, facts against a known answer. Have a person review the rest. If a change makes the set worse, it doesn’t ship.
Record the prompt and model version on every output
Store, with every output the system produces, the prompt version and the exact model version that produced it. When someone questions an answer months later, you can reproduce it, explain it, and find every other output made by the same combination. Without that record, “why did it say that?” has no answer.
Make rollback one step
Because prompts are versioned and deployed like code, a bad change rolls back like code: one revert, one deploy, back to the last version that passed. If the prompt lives in a dashboard someone edited by hand, rollback means trying to remember what it used to say.
Separate instructions from data
Prompt injection happens when data gets read as instructions: a customer email, a web page or a document that says “ignore your previous instructions”. Keep your instructions in the system prompt, put untrusted content in clearly marked sections, tell the model to treat it as data only, and never rely on the prompt alone for safety. Permissions, tool allow-lists and output validation belong in code, where no text can override them.
Document the intent
Next to each prompt, write down what it is for, what it must never do, which inputs it expects and which parts of its output the code depends on. A line like “must always return the supporting quote; the review screen depends on it” saves the next person from a well-meant edit that breaks production.
Prompt checklist
- Prompts as files in version control
- Templates with named, checked variables
- Every change reviewed in a pull request
- A fixed example set run before every release
- Prompt and model version stored on every output
- One-step rollback
- Instructions kept apart from untrusted data
- Intent and constraints documented next to the prompt
None of this is special to AI. It is the same engineering discipline we apply to every part of a system; see how we work.
Written from our engineers’ work on production systems. Want a second opinion on your project? Talk to an engineer.
See the work →