Put one gateway between your product and the AI models
When every feature calls AI providers directly, keys, costs and personal data end up scattered across the codebase. One internal gateway gives you a single place to control them, and lets you change models without touching a feature.
The first AI feature usually calls a provider directly: an API key in an environment variable, a client library, a request. The second feature copies the pattern. By the fifth, there are several keys, two providers, three ways of handling errors, and nobody can say what the AI spend is per feature or whether customer data is reaching a model it should not.
If you lead the team: what to ask
- If we need to switch AI provider tomorrow, how many places in the code change?
- Can we say what each AI feature costs us, and which customers drive that spend?
- Where are the AI provider keys kept, and is personal data masked before it’s logged or sent out?
The fix is an architectural decision that is cheap early and expensive late: every model call goes through one internal gateway. Features ask the gateway for a capability; the gateway decides how to fulfil it. You gain one place to control security, cost, reliability and model choice.
What the gateway does
- Keys and secrets in one place. Provider credentials live only in the gateway. Features never see them, rotation happens once, and a leaked feature cannot leak a key it never had.
- Routing by task. Features request a task, such as “summarise”, “classify” or “extract fields”, not a specific model. The gateway maps each task to the model that fits its quality, speed and cost needs, and that mapping lives in configuration.
- Fallback between providers. When a provider is slow, rate-limited or down, the gateway can retry on an alternative model that has been tested for that task. The feature gets an answer, along with a record of which model produced it.
- Rate limits. Limits per feature and per customer protect your provider quotas, so one heavy user or one runaway job cannot starve everything else.
- Cost tracking per feature and customer. Every call is tagged with the feature and the customer it serves, with tokens and cost recorded. You can answer “what does this feature cost us?” and “which customers drive our AI spend?” from data, and set budgets and alerts on top.
- Logging and redaction. Requests and responses are logged in one consistent format for debugging and audit. Personal data is detected and masked before it is written to logs, and where policy requires it, before it is sent to an external model at all.
- Caching. Identical requests, such as the same document classified twice, can be answered from a cache, which saves time and money. Cache rules belong in the gateway, where they can be applied consistently and switched off for anything that must always be fresh.
- Swapping models without touching features. When a better or cheaper model appears, you change the routing, run the task’s tests, and roll it out. No feature code changes.
How to start
A gateway does not need to be a large platform. It can begin as a single internal module or service with one function per task and a configuration file for routing. What matters is the rule: no feature calls a model provider directly. Enforce it in code review, or by giving only the gateway network access and credentials for the providers.
Then add capabilities in order of need. Most teams start with centralised keys, logging and cost tagging, because those answer the first questions management asks. Fallback and routing follow when a second provider or model is in use. Redaction comes first, not later, wherever personal or regulated data is involved.
What to watch for
- A fallback is a different model. Test each fallback against the same cases as the primary model before relying on it, and record which one answered.
- The gateway is now critical. Give it the same care as any shared service: health checks, timeouts, monitoring and a deployment process.
- Keep features honest. Validation of each feature’s output still belongs to the feature. The gateway handles transport and policy; it does not make an answer correct.
If you want an outside view of where AI calls, keys and data flow today in your product, our AI readiness check is a practical place to start.
Gateway checklist
- Does every model call go through one internal layer?
- Are provider keys held only by that layer?
- Do features request tasks, with model choice in configuration?
- Is there a tested fallback for each critical task?
- Are rate limits set per feature and per customer?
- Is every call tagged with feature, customer, tokens and cost?
- Is personal data masked in logs and, where required, before sending?
- Can you change a model without editing feature code?
Written from our engineers’ work on production systems. Want a second opinion on your project? Talk to an engineer.
See the work →