Data contracts: agree on the shape before you connect two systems
Integrations break quietly when two systems disagree about what a record looks like. A data contract makes that agreement explicit, testable and owned, so bad data stops at the boundary instead of reaching your reports and your AI.
Most integration failures are not outages. The connection works, the data flows, and somewhere downstream a report is wrong, a customer appears twice or an AI assistant answers with a date that is off by a day. The two systems never disagreed about whether to talk. They disagreed about what the data means.
If you lead the team: what to ask
- Is there a written agreement on what each field means, and who owns it?
- If the other system sends a record that doesn’t fit, is it stopped with a clear error, or quietly let through?
- Before either side changes a field, who has to be told, and do our tests catch the change?
A data contract settles that before the first record moves. It is a written, versioned agreement between the system that produces data and the systems that consume it, and it is enforced in code. You gain integrations that fail on day one, loudly and in one place, instead of drifting for months.
What a data contract covers
A useful contract answers the questions that integrations usually leave to guesswork:
- Schema. Every field, its type and its format. “Amount” is a decimal in a stated currency, not a string that sometimes contains a comma.
- Required and optional fields. Which fields must always be present, which may be empty, and what empty means: unknown, not applicable or zero are three different things.
- Identifiers. Which ID is the stable key for each record, who issues it, and whether it can ever change. Email addresses and names are not IDs.
- Timestamps and time zones. Every date carries its time zone, or the contract fixes one (UTC is the usual choice). It also says what each timestamp records: when the event happened, when it was entered, or when it was last modified.
- Enumerations. The full list of allowed values for statuses, types and stages, and what a consumer must do when it meets a value it does not know.
- Versioning. How changes are numbered, which changes are compatible (adding an optional field) and which are breaking (renaming, removing or changing the meaning of a field).
- Ownership. Who owns the contract, who must be told before a change, and how much notice consumers get.
The format matters less than the habit. JSON Schema, OpenAPI, typed models in the codebase or a well-kept table all work, as long as the contract lives next to the code and both sides can read it.
Contract tests: the agreement, checked on every change
A contract that only lives in a document will be broken by the next sprint. Contract tests turn it into something a pipeline can check. The producer’s tests confirm that what it emits matches the contract. The consumer’s tests confirm that it handles every valid shape, including optional fields left empty and enum values it has never seen. When either side changes, the build fails before the change reaches production, which is the cheapest moment to find out.
Fail loudly at the boundary
Every integration has a boundary: the point where outside data enters your system. That is where the contract should be enforced. A record that does not match is rejected or quarantined there, with a clear error that names the source, the record and the field.
The alternative is quiet acceptance. A missing ID gets a default, an unknown status gets mapped to “other”, a date without a time zone gets read as local time. Each choice looks harmless. Together they spread errors into every report, dashboard and model that reads the data, where they are far harder to trace. One loud rejection at the edge costs minutes; a quiet error found weeks later costs an investigation.
Why AI depends on it
An AI assistant or scoring model is only as reliable as the context you give it. If “closed” means won in one system and lost in another, the model will reason confidently over a contradiction. If timestamps mix time zones, its answer to “what happened yesterday” is wrong for part of the data. Models do not flag these problems; they smooth over them.
Data contracts are the cheapest way to make AI context consistent: the shape, the meaning and the identity of every record are settled before the model ever reads it. It is the same normalization work we describe in AI on your data, made explicit and enforceable.
Checklist before you connect two systems
- Is every field typed, with required and optional fields stated?
- Is there one stable ID per record, and do both sides agree on it?
- Does every timestamp carry a time zone and a defined meaning?
- Are enum values listed, with a rule for unknown values?
- Is there a version number and a rule for breaking changes?
- Do contract tests run on both sides in CI?
- Does invalid data stop at the boundary with a clear error?
- Is there a named owner who approves changes?
Written from our engineers’ work on production systems. Want a second opinion on your project? Talk to an engineer.
See the work →