Ingestion, normalization, trust: feeding practitioner data to an AI assistant
An AI assistant for legal and healthcare practitioners is only as good as the records it can read. What we learned owning the integration and data layer of such a platform, and why calling an endpoint is the easy part.
An AI assistant for lawyers and clinicians sounds like a model problem. In practice, much of it is a data problem. The assistant needs each practice’s own records: clients and patients, notes, test results, cases. Those records live in systems that were never designed to talk to each other, let alone to an AI.
We have done exactly that work on a production AI platform for legal and healthcare practitioners: API integration, data ingestion and normalization, inside a larger product built by a wider team.
Calling the endpoint is the easy part
The sources included legal case-management systems and healthcare practice systems, each with its own idea of what a “contact”, a “note” or a “result” is. They spoke different dialects too: REST here, SOAP and XML there, OAuth2 in one place, API keys or webhooks in another.
The real problem was never reaching the data. It was making incompatible source data trustworthy and usable by everything downstream.
One adapter per source, one model for everyone
- Adapters speak each external system’s language, its authentication and quirks included, and nothing else.
- Typed, validated models sit at the boundary. Practices, contacts and patients, notes and results each have one shape, checked on the way in, so a malformed record fails loudly instead of quietly polluting the assistant’s context.
- Repositories and services above them never need to know which system a record came from.
Assume the other side will fail
External legal and health systems throttle, time out, return inconsistent data and change without notice. The integration layer was built for that from the start:
- retries and timeouts on every call;
- circuit-breaker and rate-limit patterns, so one failing source cannot take the rest down;
- structured logs with correlation IDs, so a single record can be followed from source to assistant.
Regulated data changes the bar
Client files and patient records are among the most sensitive data there is. Access control, credential handling, separation between practices and traceable data movement were engineering requirements, not paperwork, under SOC 2, HIPAA and GDPR expectations.
Five stages
The work left us with a simple model for any AI system that runs on someone else’s data:
Ingestion → Normalization → Trust → Value → Production
The model can only add value once the first three are true. Most of the effort, and most of the risk, lives there.
Where our work sat
We owned substantial integration, ingestion and normalization work. The assistant itself, its agents, workflow orchestration, search and retrieval were the platform team’s, and we worked within them. Our part was making the data they all depend on fit to be trusted.
This article is drawn from a real engagement. Client details withheld; every figure comes from the client’s own data.
Read the full case study →