Blog / Platforms & migrations

Migrating the data is the real migration

New code can be rewritten as often as needed. Years of business records cannot. How to map, clean, rehearse, reconcile and cut over the data when you replace an old system, with a way back and the old records still readable.

When teams plan the replacement of an old system, most of the attention goes to the new code: the stack, the screens, the architecture. The code is the part you can redo. The data is not. Years of customers, orders, invoices and history are what the business actually runs on, and if they arrive wrong, the new system is wrong from its first day, however good its code. Migrating the data is the real migration; the rest is software.

If you lead the team: what to ask

  • How many times have we rehearsed the full data migration on a copy, and did the numbers match each time?
  • After the switch, how will we prove that every record arrived, and arrived correctly?
  • If we have to go back, what happens to the records created in the new system in the meantime?

Map every field, explicitly

Start with a written mapping: for each table and field in the old system, where it goes in the new one, how it is transformed, or why it is left behind.

  • Types and formats. Dates stored as text, amounts with implied decimals, time zones nobody recorded, character encodings from another era.
  • Meanings. A status code means what the old code did with it, not what its name says. Confirm each value with the people who use it.
  • Fields with no home. Decide deliberately.

Clean before you move, and know what you changed

Old data carries old problems: empty required fields, values outside any valid range, references to records that no longer exist. Some can be fixed by rules, some need a person to decide, some must be carried over as they are. Whatever the choice, the cleaning is done by scripts, never by hand, and every change is logged so anyone can see what was altered and why.

Handle IDs and duplicates on purpose

Identifiers are where migrations quietly break. Other systems, old emails, printed invoices and bookmarks refer to records by their old IDs. Keep a mapping table from every old ID to its new one, and keep the old ID as a field on the new record.

Duplicates need rules agreed with the business: when two customer records are the same customer, which values win, and where do their orders, notes and history go. Record every merge.

Rehearse on copies until it is boring

The migration is a script, or a set of them, and it runs many times against fresh copies of production data before it runs for real. Each dry run produces the same report. When two runs in a row produce the same clean numbers, and the run time fits the window you have, the rehearsals have done their job.

Reconcile with counts and checksums

“It looks right” is not evidence. Reconciliation is:

  • Counts per table and per category: customers, active customers, orders per year, invoices per status. Every difference explained.
  • Totals on the numbers that matter: balances, invoice amounts, quantities. These must match to the cent.
  • Checksums on key fields, computed the same way on both sides, to catch values that changed silently.
  • Samples checked by people who know the data, on the records they care about most.

The same checks run after the real migration, and their results are kept.

Plan the cut-over, and the way back

Write the cut-over as a timed runbook: freeze or pause writes to the old system, final migration run, reconciliation, switch, smoke tests, and the moment of decision. Name who decides, and what result means go or stop.

Then plan the rollback before you need it. Going back in the first hour is simple. Going back after users have created records in the new system is not: decide in advance whether those records will be copied back, re-entered or held, and how. If you replace the system piece by piece, each piece needs its own version of this plan.

Keep the old data readable

Once the switch is done, keep a read-only copy of the old data, and a way to query it, for as long as legal, audit or support needs require. Questions about old records arrive for years, and the answer is often a field that was never migrated.

Where we come from

We have done this kind of work on a production AI platform for legal and health practitioners, where we owned the ingestion and normalization of records from many systems into one validated model. The lesson carries over directly: data has to be made trustworthy on the way in, or everything downstream inherits its problems. It is also why data gets as much care as code in our legacy modernization work.

Data migration checklist

  • Is there a written mapping for every field, including what is left behind and why?
  • Is cleaning done by logged scripts, with each change traceable?
  • Does every new record keep its old ID, with a full mapping table?
  • Are duplicate rules agreed with the business, and every merge recorded?
  • Has the full migration run on fresh copies until the numbers are stable?
  • Do counts, totals and checksums reconcile, and are the results kept?
  • Is there a timed cut-over runbook with a named decision-maker?
  • Does the rollback plan cover records created after the switch?
  • Will the old data stay readable for as long as it may be needed?

Written from our engineers’ work on production systems. Want a second opinion on your project? Talk to an engineer.

See the work →

Want us to look
at your site?

Tell us where traffic, revenue or your numbers stopped making sense. We will tell you what we would check first.

Book a 15-min call

We only send what you ask for. Privacy

Prefer to write directly? enable JavaScript to see the address

Talk to an engineer

No sales theater. Tell us where your operation feels slow, repetitive or difficult. An engineer reads every message and replies by email.

Prefer to talk? Pick a 15-minute slot →

Your message goes straight to our engineers at our address.