Agent or workflow? When to let the model decide
Most AI features are more reliable, cheaper and easier to test as a fixed workflow with AI in a few steps. How to tell when an autonomous agent is worth its cost, with a checklist you can use on your next feature.
“Agent” has become the default word for any AI feature. But there is a real design choice underneath it, and getting it right decides how predictable, testable and affordable your system will be. The question is simple: who decides the next step, your code or the model?
If you lead the team: what to ask
- Could we write the steps of this feature down in advance? If so, why is the model choosing them?
- Does the same input need to give the same result every time?
- If the AI takes a wrong step, can it be undone, and does a person see it first?
Two shapes for an AI system
In a workflow, your code fixes the sequence. Step one, step two, step three, every time. Some steps call a model to do what models are good at, such as writing, classifying or extracting, but the model never chooses what happens next. The path is known before the run starts.
In an agent, the model drives. It receives a goal and a set of tools, decides which tool to call, reads the result, and decides again, until it judges the goal met. The path is discovered during the run.
Both are legitimate. They solve different problems, and they carry very different costs.
A workflow with AI where it adds value
The content pipeline behind SuperEscuela, the free learning platform we built, has four stages, and AI works in only two of them:
- Plan (AI). The model plans the topic tree: what comes first and what each topic should teach.
- Walk the plan (plain code). A script goes through the tree and sends each topic on, in order.
- Write (AI). The model writes the lessons for each topic.
- Publish (plain code). Code exports and publishes the result.
Walking a list and publishing files are predictable steps, and predictable steps should behave the same way every time. Giving them to a model would add variation where none is wanted. The AI does the creative work; the code does the bookkeeping.
When an agent earns its place
An agent is worth considering when the path genuinely cannot be written down in advance:
- The task is open-ended. Investigating an issue, researching a question or exploring a codebase, where each finding changes what to look at next.
- Tool choice depends on the content. There are many possible tools and which ones are needed, and in what order, only becomes clear while working.
- The number of steps varies widely. Some requests need one action, others need twenty.
- A person reviews the outcome. The result is a draft, a proposal or a finding, not an irreversible action.
What autonomy costs
- Unpredictability. The same input can take a different path on each run. That is the feature, and also the problem when you need consistent results.
- Testing. A workflow can be tested step by step against known answers. An agent has to be evaluated across many runs and many paths, which takes more effort to design and to maintain.
- Blast radius. An agent can do anything its tools allow. Every tool you hand it widens what can go wrong, so permissions, limits and confirmations become part of the design.
- Cost and latency. Each decision is another model call. Loops that a workflow would finish in three calls can take many more.
- Debugging. When a workflow fails, you know which step failed. When an agent fails, you need a full trace of what it decided and why.
The usual answer is a mix
In practice the strongest designs are workflows by default, with model calls in the steps that need judgment, and an agent confined to the one part that is genuinely open-ended. That agent gets a narrow set of tools, a step limit, a timeout, and its output passes through the same validation as everything else. Deciding which parts belong to code, to AI and to people is a core step of how we work.
Agent or workflow: a decision checklist
- Can you write the steps down in advance? If yes, build a workflow.
- Must the same input give the same result? Prefer a workflow.
- Does choosing the next step really require judgment about content? An agent may fit.
- Can its actions be undone, or does a person review them first? If not, keep it out of an agent’s reach.
- Can you limit its tools, steps, time and spend? If not, it is not ready.
- Do you have a way to evaluate many runs, not just one demo?
- Will every decision be logged so a failure can be traced?
This article is drawn from a real engagement. Client details withheld; every figure comes from the client’s own data.
Read the full case study →