Ingestion: start where the business actually runs
The first stage of every project is getting the real data in, from where the business actually operates. Why a sample is not enough, and what a traffic diagnosis and a sales product taught us about it.
The fastest way to a wrong conclusion is a clean sample. It is tidy, it fits in a spreadsheet, and it leaves out the part of the business where the problem lives.
So the first stage of every project we take on is ingestion: getting the business’s real data into one place, from the systems where it actually operates.
A traffic collapse, explained by six sources
When a content and affiliate site lost most of its search traffic, the easy move was to open the analytics dashboard and start guessing. Instead we brought in:
- Search data: 16 months of it, by page and by search phrase, not the few months the interface shows.
- Code history: every release, plugin update and settings change, with its date.
- Server logs: what the server actually received, not what the analytics tag reported.
- Site snapshots and speed measurements: how pages looked and performed over time.
- A read-only database account: what the site really contained.
- Revenue history: because traffic is not the goal; money is.
Only with all of them side by side could each step of the decline be dated and tied to a cause. And one finding was invisible from any single source: more than half of the visits the analytics reported were bots. The dashboard alone could never have shown that.
A sales product, fed from everywhere
An AI sales-coaching product needs to see what the sales team sees. That information was scattered: the CRM, call recordings and transcripts, documents, spreadsheets, chat and email. Ingesting it meant one connection per source, built around each one’s format and limits. The CRM connection only reads, so the product learns from the customer’s pipeline without ever changing it.
What good ingestion looks like
- All of it, not a sample. Problems hide in the second pipeline, the old pages, the quiet months.
- Read-only by default. Diagnosis should never risk the system being diagnosed.
- Every record traceable. Know where each piece of data came from, and when.
- Built for the source’s bad days. APIs throttle, exports break, logs rotate. Ingestion has to cope.
Getting the data in is the start, not the finish. Six sources means six ways of describing the same thing, which is the next stage: normalization.
This article is drawn from a real engagement. Client details withheld; every figure comes from the client’s own data.
Read the full case study →