Blog / Custom software

Characterization tests: write down what the old system does before you change it

Before you refactor or replace an old system, record what it actually does today, odd parts included. How characterization tests, golden masters and captured real traffic give you a safety net, and how AI speeds up writing them.

The riskiest moment in a legacy project is the first change. There are no tests, no current specification, and the people who knew why it works that way have moved on. Before you change anything, you need a record of what the system does today. That record is a set of characterization tests.

If you lead the team: what to ask

  • Before we change the old system, do we have tests that record what it does today, including the strange parts?
  • Were those tests built from real inputs, or only from cases someone imagined?
  • When the new code behaves differently from the old, who decides whether that is a fix or a regression?

A normal test checks that the code does what it should. A characterization test checks that the code does what it does. It doesn’t judge. It describes. If the old invoice module rounds a total down instead of to the nearest cent, the characterization test expects the rounding down. Whether that is correct is decided later, on purpose, not by accident in production.

Why “what it does” beats “what it should do”

An old system in daily use has been shaped by years of fixes, exceptions and workarounds. Customers, reports and other systems depend on its exact behaviour, including behaviour nobody would design today. If you write tests from the specification, you test your idea of the system. If you write them from observation, you test the system your business actually runs on.

The goal is a safety net: once current behaviour is pinned down, any change that alters it fails a test.

Golden masters and snapshots

Writing one assertion per behaviour is slow on large, unfamiliar code. The faster approach is the golden master:

  1. Choose an entry point: a function, an endpoint, a batch job, a report.
  2. Feed it a broad set of inputs.
  3. Record every output in full, as files: the response body, the generated document, the rows written to the database.
  4. From then on, run the same inputs and compare the new outputs with the recorded ones. Any difference fails the test.

Snapshot testing is the same idea at a smaller scale: the first run saves the output, later runs compare against it. Both need care. Strip or freeze anything that changes on its own, such as timestamps, generated IDs, random values and the ordering of unordered collections, or the tests will fail for reasons that have nothing to do with your change.

Capture real inputs and outputs

The best inputs come from production, because production contains the cases nobody thought of. Useful sources:

  • Request logs for APIs and web endpoints, replayed against a test copy.
  • Input files that batch jobs have processed, with the files they produced.
  • A recording layer added temporarily in front of the old code, saving each request and response as it happens.

Two rules apply. Personal and sensitive data is masked or replaced before it goes anywhere near a test suite. And sampling is deliberate: include the common cases, but also the edges, such as zero amounts, very long names, old record formats and dates at month end and year end, because the rare cases are where old systems hide their rules.

Where AI speeds this up

This is work AI tools do well, with a person checking the result:

  • Reading the code. Ask the model to list every branch in a function and the input that would reach each one.
  • Drafting tests from traffic. Give it a set of captured requests and responses and the module’s code, and have it write the tests: setup, inputs, expected outputs.
  • Filling the gaps. Run the tests with a coverage tool, show the model the lines never executed, and ask for inputs that would reach them.

The rule: the expected values come from running the real old system, never from the model’s guess about what the code does. The model writes the test; the old system supplies the answer.

Decide later which odd behaviours are bugs

A characterization suite will pin down behaviour that looks wrong. Don’t fix it while you are building the net. Tag those tests, list them, and review the list with the people who own the business process. Each one becomes one of three decisions: keep it (someone depends on it), fix it (a real bug, and the test changes along with the fix, on purpose), or retire it (nobody uses that path any more).

This is how we start our legacy modernization work: record first, change second.

Checklist: before the first change

  • Entry points chosen: the functions, endpoints, jobs and reports the change will touch.
  • Real inputs captured, with personal data masked.
  • Edge cases sampled on purpose: zero, empty, very long, old formats, month and year end.
  • Outputs recorded in full from the old system as the golden master.
  • Timestamps, IDs and random values frozen or removed from comparisons.
  • AI-drafted tests reviewed, with expected values taken from the old system, not the model.
  • Coverage checked, and gaps filled with targeted inputs.
  • Odd behaviours listed and sent to the business owners: keep, fix or retire.

Written from our engineers’ work on production systems. Want a second opinion on your project? Talk to an engineer.

See the work →

Want us to look
at your site?

Tell us where traffic, revenue or your numbers stopped making sense. We will tell you what we would check first.

Book a 15-min call

We only send what you ask for. Privacy

Prefer to write directly? enable JavaScript to see the address

Talk to an engineer

No sales theater. Tell us where your operation feels slow, repetitive or difficult. An engineer reads every message and replies by email.

Prefer to talk? Pick a 15-minute slot →

Your message goes straight to our engineers at our address.