Blog / Platforms & migrations

Limit the blast radius: decide what a failure is allowed to break

Every integration and every AI feature will fail at some point. Deciding in advance how far that failure can spread is what keeps one bad source, one bad deploy or one bad model output from taking the rest of the business with it.

Things will break. An external API will time out, a vendor will change a field without notice, a deploy will carry a bug, a model will produce something nobody expected. You cannot prevent all of that. What you can decide, ahead of time, is how much each failure is allowed to break. That is the blast radius, and it is a design decision, not luck.

If you lead the team: what to ask

  • If one connected system goes down, what else stops working?
  • Can we switch off a new AI feature in minutes, without a deploy?
  • What is the AI allowed to do on its own, and is that limit enforced in code rather than in the prompt?

What it looked like on a practitioner data layer

On the data layer of an AI platform for legal and health practitioners, we connected to many external systems that throttle, time out and return inconsistent data. Every external call had retries and timeouts. Circuit-breaker and rate-limit patterns meant one failing source could not take the others down: if one system stopped responding, calls to it were cut off quickly, and every other source kept flowing. Structured logs with correlation IDs made it possible to follow a single record and see exactly where a failure stopped.

The same thinking applies to any system that depends on integrations or AI. Here are the tools we reach for.

Isolate per source and per tenant

A failure in one source, or in one customer’s data, should stay there. In practice:

  • separate queues, workers or connection pools per external source, so one slow system cannot exhaust shared resources;
  • a circuit breaker per source, not one for the whole integration layer;
  • tenant separation enforced in the data layer, so a bug in one customer’s processing cannot read or write another’s records;
  • bad records quarantined for review instead of halting the whole import.

Read-only by default, scoped credentials always

The cheapest way to limit damage is to make it impossible. If an integration only needs to read, give it read-only access. If it needs to write, scope the credential to the objects and actions it actually uses. One credential per integration, not a shared admin key, means a leaked or misused key opens one door instead of all of them, and can be revoked without stopping everything else.

Decide what an AI action may touch

AI features raise the question sharply. A model that suggests is low risk; a model that acts is not. For every AI capability, write down:

  • which data it can read, and which it must never see;
  • which actions it can take on its own, which need a person’s approval, and which are off limits;
  • the most it can do in one run: records changed, messages sent, money moved.

Enforce those limits in code around the model, not in the prompt. A prompt is a request; a permission check is a rule.

Feature flags and kill switches

Every new integration and every AI feature should ship behind a flag that someone can turn off without a deploy. A kill switch is the fastest way to shrink a blast radius that is already growing: switch the feature off, the rest of the product keeps working, and you investigate calmly.

Fallbacks that degrade gracefully

When a dependency fails, decide what the user sees. Cached data with a clear “last updated” note, a rule-based answer instead of the model’s, a queued request that completes later. A fallback turns an outage into a slower or simpler experience instead of an error page.

Staged rollouts

Release to a small group first: internal users, one tenant, a share of traffic. Watch errors, fallbacks and corrections, then widen. A problem found at that stage affects a few people instead of everyone, and rolling back is routine instead of an emergency.

Blast-radius checklist

  • Timeouts and bounded retries on every external call.
  • A circuit breaker and rate limit per source.
  • Tenant separation enforced below the application code.
  • Read-only by default; one scoped credential per integration.
  • A written list of what each AI capability may read and do.
  • A flag and a kill switch for every new feature.
  • A defined fallback for each critical dependency.
  • Staged rollout, with correlation IDs to trace what went wrong.

This is part of how we design every system; see our method for the rest.

This article is drawn from a real engagement. Client details withheld; every figure comes from the client’s own data.

Read the full case study →

Want us to look
at your site?

Tell us where traffic, revenue or your numbers stopped making sense. We will tell you what we would check first.

Prefer to write directly? enable JavaScript to see the address

Talk to an engineer

No sales theater. Tell us where your operation feels slow, repetitive, or difficult — an engineer reads every message.

Your message goes straight to our engineers at our address.