AI, in plain words

Blog / AI systems

AI risks in plain language: what can go wrong, and how to limit it

Six real ways an AI feature goes wrong, from confident wrong answers to surprise bills: what each looks like in practice, and the plain controls that keep each one small.

AI features fail differently from ordinary software. A normal program either works or shows an error. An AI system can fail while looking perfectly fine: a polished answer, a green check mark, and a wrong result underneath. That is why the risks are worth naming one by one. None of them is a reason to avoid AI. Each has a known way to keep it small.

1. Confident wrong answers

What it looks like: the assistant quotes a policy that doesn’t exist, a summary invents a figure, a report describes a trend the data doesn’t show. It reads just as confidently as a correct answer.

How to limit it: make the AI show its sources, so a person can check them. Have plain code verify whatever can be verified: totals add up, dates are real, every required field is present. Test the system on questions where you already know the answer, before and after every change. And where an answer leads to a decision, a person reviews it first.

2. Leaking data

What it looks like: an employee pastes customer records into a public chatbot. An internal assistant answers one client’s question with another client’s information. A log file quietly keeps every prompt, personal details included.

How to limit it: decide which tools are approved and what data may go into them. Use providers whose terms say your data isn’t used to train their models. Keep each customer’s data separate in the system itself, not by asking the AI to be careful, and test that separation. Keep sensitive details out of logs.

3. Doing something it shouldn’t

What it looks like: an AI that can act, not just answer, sends an email to the wrong list, deletes records it was asked to tidy up, or follows instructions hidden inside a document it was reading.

How to limit it: give it the fewest permissions that get the job done, like a new employee who only gets keys to the rooms they need. Let it read freely, but let it change things carefully. Anything that can’t be undone (sending, paying, deleting) waits for a person to approve it. Keep a record of every action it takes.

4. Cost surprises

What it looks like: AI providers charge per use. A loop that retries forever, a feature that sends whole documents when a paragraph would do, or a bot hammering your site can turn a small monthly bill into a large one overnight.

How to limit it: set spending limits and alerts with the provider. Cap how much each user and each feature can use per day. Measure the cost per task (per answer, per document) so you know what normal looks like, and you notice when it changes.

5. Vendor changes

What it looks like: the provider retires the model you built on, changes its price, or updates it so that answers that used to be right come out slightly different. Nothing in your code changed, but your results did.

How to limit it: use the exact model version you tested, not “the latest.” Keep your known-answer tests and run them before switching versions. Build so the provider can be swapped without rewriting the system. Read the provider’s retirement notices; they usually arrive months ahead.

6. Nobody knowing how it works

What it looks like: the person who set it up leaves. The instructions given to the AI (the prompts) live in someone’s chat history. When results get worse, nobody knows what changed or how to fix it.

How to limit it: treat prompts like code: stored in one place, versioned, reviewed before they change. Write down what the system does, what data it uses, which model, and who owns it. One page is enough, as long as it stays current. Name one person who is accountable for it.

How big is each risk for you?

Not every feature needs every control. A tool that drafts internal notes for a person to edit carries little risk. A system that answers customers, touches money or acts on its own carries much more. Ask two questions for each feature: what is the worst realistic mistake, and who would notice it? The worse the mistake and the later it would be noticed, the more checks and approvals it needs.

The common thread

Every control above is ordinary engineering: checks, permissions, approvals, limits and documentation. AI brings the speed. People and plain software keep it honest. If you want to see how we put these in place step by step, our method lays it out.

Risk checklist for any AI feature

  • Answers show their sources, and plain code checks what can be checked.
  • A set of known-answer tests runs before every change.
  • Approved tools and allowed data are written down.
  • The system keeps customer data apart, and that separation is tested.
  • The AI has only the permissions it needs; actions that can’t be undone need a person’s approval.
  • Spending limits and alerts are on, and you know the cost per task.
  • The model version is fixed, and you can switch providers.
  • Prompts are stored and reviewed like code, and the feature has a named owner.

Written from our engineers’ work on production systems. Want a second opinion on your project? Talk to an engineer.

See the work →

Want us to look
at your site?

Tell us where traffic, revenue or your numbers stopped making sense. We will tell you what we would check first.

Prefer to write directly? enable JavaScript to see the address

Talk to an engineer

No sales theater. Tell us where your operation feels slow, repetitive or difficult. An engineer reads every message and replies by email.

Prefer to talk? Pick a 15-minute slot →

Your message goes straight to our engineers at our address.