Guardrails: decide what the AI must never do
Guardrails turn “the model should not do that” into rules the system enforces. How to set them on what goes in, what comes out and what the AI is allowed to do, and how to test and maintain them like code.
Most AI projects spend their energy on what the model should do. The decisions that protect the business are about what it must never do: answer outside its scope, leak someone’s personal data, invent a policy, promise a discount, send a message nobody approved. Writing “please don’t” in the prompt is not enough. A prompt is guidance the model usually follows. A guardrail is a rule the system enforces whether the model cooperates or not.
If you lead the team: what to ask
- What is this AI feature not allowed to say or do, and where is that list written?
- If a document or email contains hidden instructions, can it change what the AI does?
- Have we tried to break the guardrails on purpose, and do those tests run on every change?
Guardrails belong in three places: on what goes in, on what comes out, and on what the AI is allowed to do.
Input guardrails
- Scope. Decide what the feature is for and turn away the rest early. A support assistant for your product does not need to write essays or give medical advice. A simple classifier or rule set in front of the model can route out-of-scope requests to a polite refusal or a person.
- Prompt injection. Any text the model reads, from a user, an email, a document or a web page, can contain instructions. Keep system instructions separate from user and retrieved content, mark untrusted content clearly, and never let text from a document change what the system is permitted to do.
- Personal data. Detect and mask personal data before it reaches the model when the task does not need it, and before it reaches your logs in every case.
Output guardrails
- Schema. Require structured output and validate it. A response with missing fields or the wrong types is rejected or retried, never shown.
- Banned claims. List what the AI must never state: prices or discounts it was not given, guarantees, legal or medical conclusions, statements about competitors. Check outputs against that list with rules, and with a second model where wording varies.
- Citations required. When the answer must come from your documents or data, require the output to point to its sources, and verify that the cited sources exist and contain what is claimed. No source, no answer.
- Refusal rules. Define when the system says “I can’t answer that” or hands over to a person, and make that response clear and useful rather than evasive.
Action guardrails
When AI can act, through tools, APIs or automations, the guardrails that matter most are around actions:
- Allow-lists. The model can call only the tools and endpoints you list, with parameters validated in code. Everything else is denied by default.
- Approval for consequential actions. Sending to customers, changing records of truth, moving money or deleting anything goes through a person or a strict rule check first.
- Rate and volume limits. Cap how many actions an AI can take per run, per user and per hour, so a loop or a misunderstanding cannot repeat itself thousands of times.
Test guardrails with adversarial examples
A guardrail you have not tried to break is an assumption. Build a test set of inputs designed to get past it: out-of-scope requests phrased politely, instructions hidden inside documents, requests for personal data, attempts to get a banned claim through rewording, tool calls with out-of-range parameters. Run that set on every change to prompts, models or rules, and add every real incident to it. Track two things: what got through, and what was blocked that should not have been, because over-blocking also costs users.
Guardrails are code
Keep guardrails in the repository, versioned, reviewed and tested like any other code. A banned-claims list edited in a spreadsheet, or a filter changed directly in production, is a rule nobody can audit. Each guardrail should have an owner, a reason it exists, tests that prove it works, and a log entry every time it fires, so you can see which rules are doing real work and which are noise.
Guardrails checklist
- A written scope, with out-of-scope requests routed before the model.
- System instructions separated from user and retrieved content.
- Personal data masked before the model when not needed, and always before logs.
- Schema validation on every output.
- A banned-claims list, checked automatically.
- Citations required and verified where answers must come from your data.
- Tool allow-lists, approvals for consequential actions, rate limits.
- An adversarial test set run on every change.
- Guardrails versioned, reviewed and logged when they fire.
To see where your AI plans stand on these, take the AI readiness check.
Written from our engineers’ work on production systems. Want a second opinion on your project? Talk to an engineer.
See the work →