Blog / Platforms & migrations

Replace it piece by piece while it keeps running: the strangler fig pattern

A full rewrite of old software asks the business to stop and wait. The strangler fig pattern replaces a legacy system one capability at a time, with old and new running side by side and a way back at every step.

The old system works, mostly. It also slows every change, depends on people who know its corners, and runs on technology fewer engineers want to touch. The tempting answer is a full rewrite: build the new system beside it, then switch over one weekend. That plan asks the business to wait months for value and puts every unknown into a single cut-over. There is a safer way: replace the old system piece by piece, while it keeps running.

If you lead the team: what to ask

  • Which part of the old system will we replace first, and why that one?
  • While old and new run side by side, how will we know they give the same results?
  • If the new part misbehaves in production, how fast can we send traffic back to the old one?

The approach is known as the strangler fig pattern, after a plant that grows around a host tree until it can stand on its own. The new system grows around the old one, takes over its work a capability at a time, and the old system is retired only when nothing depends on it any more.

Put a routing layer in front

Everything starts with a single point that all requests pass through before they reach the old system. It can be a reverse proxy, an API gateway, or a thin layer of code in the application itself.

That layer is what makes the rest possible. Once it exists, sending one kind of request somewhere else is a configuration change, not a rewrite. It is also where you add logging, so you can see which parts of the old system are actually used. Features nobody calls can often be retired without being rebuilt at all.

Move one capability at a time

Pick a capability with clear edges: a report, a search, a pricing calculation, a set of pages. A good first candidate is valuable enough to matter, small enough to finish in weeks, and loosely tied to the rest. Build it in the new system, with tests that describe what it must do.

Then point the routing layer at it, for a small share of traffic or a group of internal users first. The rest of the old system keeps running untouched.

Run old and new side by side, and compare

Before the new part takes real traffic, let both answer the same requests. The old system still serves the user; the new one runs in the background (often called shadow mode), and its result is recorded next to the old one.

  • Compare outputs automatically. Totals, statuses, records returned, rendered fields. A script that flags every difference is worth more than a manual spot check.
  • Investigate every difference. Some are bugs in the new code. Some reveal a business rule the old code applied that nobody had written down. Both are worth finding before users do.
  • Decide on the expected differences. If the old system rounded wrongly or kept a known bug, record the decision to change it, so the difference is a choice and not a surprise.

Keep a way back at every step

Each switch has a written rollback: which setting in the routing layer sends the traffic back to the old path, who can change it, and how long it takes. Keep the old part running and its data current until the new one has handled real use for a while. A rollback that depends on restoring a backup is not a rollback; it is a recovery.

Data needs the most care. If the new part writes data the old one also reads, decide who owns each record during the transition, and keep both sides consistent until the old part is gone.

Retire the old part, for real

The pattern only pays off when old code actually leaves. Once a capability has run on the new side without incident, remove the old route, delete the old code and its scheduled jobs, and update the documentation. Every retired piece makes the old system smaller and the next step easier. Skipping this leaves you with two systems to maintain instead of one.

What this looks like in practice

The same principles applied when our engineers consolidated fourteen websites into one platform: an inventory first, migration scripts run many times against copies, sites moved in waves with their own checks, a written way to switch back for each wave, and code deleted rather than ported.

AI speeds up much of this work: reading old code, drafting the new version, writing comparison scripts. Engineers still decide the order, review every change and own the switch. If you are weighing how to move off an old system, our legacy modernization service works this way.

Piece-by-piece checklist

  • Is there a routing layer in front of the old system, with logging?
  • Do we know which features are still used, and which can be retired without rebuilding?
  • Is the first capability small, valuable and loosely coupled?
  • Do old and new run side by side, with outputs compared automatically?
  • Is every difference either fixed or recorded as a decision?
  • Does each switch have a written, tested rollback?
  • Is ownership of shared data clear during the transition?
  • Is old code deleted once its replacement has proven itself?

Written from our engineers’ work on production systems. Want a second opinion on your project? Talk to an engineer.

See the work →

Want us to look
at your site?

Tell us where traffic, revenue or your numbers stopped making sense. We will tell you what we would check first.

Book a 15-min call

We only send what you ask for. Privacy

Prefer to write directly? enable JavaScript to see the address

Talk to an engineer

No sales theater. Tell us where your operation feels slow, repetitive or difficult. An engineer reads every message and replies by email.

Prefer to talk? Pick a 15-minute slot →

Your message goes straight to our engineers at our address.