Blog / Platforms & migrations

Integrating APIs that fail: retries, idempotency and backoff

Every external API will time out, throttle you or change without notice. The patterns that keep an integration correct when that happens: backoff with jitter, idempotency keys, retry budgets, dead-letter queues and reconciliation.

Every external API you integrate will fail. It will time out, throttle you, return a 500 during its own deploy, send a half-written record, or rename a field without telling anyone. The question is not whether your integration meets these failures, but whether it handles them on purpose or by accident.

If you lead the team: what to ask

  • When an outside system we depend on fails, what do our users see?
  • If a request is retried, could it create a duplicate invoice, appointment or message?
  • When a record fails for good, where does it go, and who looks at it?

Where we learned it

We owned the integration and data layer of an AI platform for legal and healthcare practitioners. It connected many external systems over REST and SOAP/XML, with OAuth2, API keys and webhooks. Those systems throttle, time out, return inconsistent data and change without notice. So every call had retries and timeouts, circuit-breaker and rate-limit patterns, and structured logs with correlation IDs, and each source sat behind its own adapter. These are the patterns that matter most.

Decide which errors to retry

Retrying everything is as wrong as retrying nothing. Sort errors before you write a single retry loop:

  • Retry: timeouts, connection resets, 429 Too Many Requests, 502, 503 and 504. These are usually temporary.
  • Don’t retry: 400, 403, 404 and 422. The request is wrong, the permission is missing or the record isn’t there; sending it again changes nothing. On a 401, refresh the token once, then stop.
  • Retry with care: a 500 or a timeout on a write. The server may have done the work before it failed. That is what idempotency is for.

Exponential backoff with jitter

When a call fails, wait before trying again, and wait longer each time, doubling the delay up to a cap. Then add jitter: a random amount on each wait. Without it, every worker that failed at the same moment retries at the same moment, and a struggling server gets hit by synchronized waves. If the API sends a Retry-After header, respect it.

Retry budgets

Backoff per call is not enough. Give each job a retry budget: a limit on attempts per call and on how much of your traffic can be retries. When the budget runs out, stop and record the failure. Unlimited retries turn a partial outage into a full one, and your integration into part of the problem.

Circuit breakers and rate limits

If a source keeps failing, stop calling it for a while. A circuit breaker opens after repeated failures, fails fast during a cooling-off period, then lets a trial request through. Pair it with a client-side rate limit that stays under the provider’s quota, so you are never the reason you get throttled. One slow source should never hold up the others.

Idempotency keys for every write

A retried write can create two invoices, two appointments or two notes. Send an idempotency key with every write: a unique ID for the operation, generated once and reused on every retry, so the provider applies it only once. If the API doesn’t support keys, store them on your side and check whether the record already exists before writing again. Design your own webhook handlers the same way: webhooks get delivered more than once, and receiving the same event twice must be harmless.

Dead-letter queues

When a record has used up its retries, don’t drop it and don’t block the queue behind it. Move it to a dead-letter queue with the error, the payload and the correlation ID. Someone can inspect it, fix the cause and replay it. Nothing disappears silently.

Reconciliation jobs

Even with all of the above, some changes get missed: a webhook that never arrived, a record edited during an outage. A scheduled reconciliation job compares your copy with the source, record counts and last-updated timestamps first, then the records that differ, and repairs the gaps. It is the safety net under the safety net.

Detect contract drift

APIs change without notice. Validate every response against the shape you expect, at the adapter boundary. A missing field, a new status value or a number that arrives as text should fail loudly and raise an alert, not flow quietly into your data. Drift caught at the door is a small fix. Drift caught by a customer is an incident.

Before you ship an integration

  • Timeouts on every call
  • A written list of retryable and non-retryable errors
  • Exponential backoff with jitter, and a retry budget
  • A circuit breaker and a client-side rate limit per source
  • Idempotency keys on every write, and handlers that tolerate duplicates
  • A dead-letter queue with replay
  • A reconciliation job on a schedule
  • Response validation that alerts on contract drift
  • Structured logs with a correlation ID on every call

We build every integration this way because it is cheaper than the first incident. It is part of how we work.

This article is drawn from a real engagement. Client details withheld; every figure comes from the client’s own data.

Read the full case study →

Want us to look
at your site?

Tell us where traffic, revenue or your numbers stopped making sense. We will tell you what we would check first.

Prefer to write directly? enable JavaScript to see the address

Talk to an engineer

No sales theater. Tell us where your operation feels slow, repetitive, or difficult — an engineer reads every message.

Your message goes straight to our engineers at our address.