Are We Overcomplicating Enterprise Agentic AI?
Govern the work before you build the workforce: start with the outcome, preserve evidence and add persistent agents only when the responsibility earns them.
Scope of this piece
Govern the work before you build a workforce around it. Add persistent agents when the responsibility—not the fashion—earns them.
View the article illustration

I did not begin with agents and eventually discover workflows. The workflows, tools and skills were there surprisingly early. What I kept doing was making the agent the centre again.
What I kept doing was giving the agent a name, a role, a team and an expanding governance dossier. Then the work would force the same question back onto the table:
What outcome are we trying to produce, and what is the smallest responsible way to produce it?
That question became harder to ignore while I was working on two apparently different things at the same time. On one side, I was developing operating models for agentic AI: identity, authority, trust, assurance, performance and lifecycle. On the other, I was turning repeatable knowledge work into a folder that a team member could use through one harness and one clear front door.
The second approach looked simple. Attach the source material. Point the harness to the README. Ask for the outcome.
But the simplicity was in the interface, not in the work. The folder still had to define what good looked like, where truth came from, which steps mattered, what evidence had to survive and who could accept the result. The prompt was short because the method was not being invented from scratch every time.
Getting an output was never the whole point. The model or agent was only one component. Repeatable outcomes depended on the system around it: the inputs, the method, the skills and tools, the controls, the evidence, the human judgement and the way the result was accepted.
This is not an argument that folders replace agents. It is an argument that we should govern the work before we build a workforce around it.
This was not a conversion story
It would be convenient to tell this as a clean progression: I started with a team of agents, discovered workflows, and eventually simplified everything into files. That is not what happened.
Reusable prompts, tools, workflows and hand-offs were present early. Named agents also helped. They made responsibilities easier to explain. They gave a system character, made a team model teachable and forced questions about authority that might otherwise have remained hidden.
The problem was not the names. It was allowing the names to become load-bearing when the execution did not require persistent identity.
I kept moving through the same cycle: compose the work, personify the components, add governance, encounter coordination or assurance cost, then return to the work contract. Each pass taught me something useful. It also made me question whether enterprise agentic AI is sometimes being designed from the organisation chart inward instead of from the outcome outward.
An analyst, reviewer and writer may be useful ways to describe three modes of work. They do not automatically need to be three persistent software workers. The separation matters only when it creates a real benefit: independent challenge, different access, a protected context, concurrent work or clear operational ownership.
That same instinct can scale beyond the individual worker. If every capability becomes a worker, and every worker becomes part of a team, it is not a large step to start designing the production system around them: pipelines, orchestration, governance and, eventually, an AI factory.
Sometimes the scale, authority and operational responsibility genuinely require that machinery. But sometimes we have simply industrialised the solution before we have properly defined the work.
What problem is the factory solving?
The phrase AI factory is doing a lot of work at the moment.
At the infrastructure layer, it has a clear meaning. NVIDIA defines an AI factory as specialised computing infrastructure for managing the AI lifecycle, from data ingestion through training, fine-tuning and high-volume inference. That is a legitimate scale problem involving energy, compute, models, software and operations.[1]
But that is different from composing an application to produce a bounded business outcome. And both are different again from operating sustained software workers as a workforce.
Those three problems can overlap, but they should not be collapsed:
- Infrastructure: How do we provide and operate AI capability at scale?
- Work composition: What combination of rules, tools, skills, workflows and judgement can produce this outcome?
- Workforce operation: How do we manage software that carries persistent, delegated or consequential responsibility?
If I am building shared enterprise AI infrastructure, factory thinking may be exactly right. If I am producing one bounded impact brief for a human decision, copying a factory line or an organisation chart into software may introduce more boundaries than the work earns.
So my starting question is no longer, “How many agents do we need?”
What is the smallest responsible composition that can perform this work?
The work contract is the stable asset
Models change. Tools change. Harnesses change. The actor applying a method can change as well.
What should remain stable is the contract for the work:
- the outcome and accountable owner;
- the sources that are allowed, required or excluded;
- the core method and mandatory gates;
- the evidence and intermediate artefacts that must survive;
- the authority to make consequential decisions; and
- the conditions under which the result can be accepted.
This is where Jake Van Clief's work has had a significant influence on my enterprise AI journey. In Interpretable Context Methodology: Folder Structure as Agentic Architecture, Van Clief and David McDermott describe a Model Workspace Protocol in which numbered folders, markdown files and local scripts coordinate sequential, human-reviewed work without requiring a multi-agent framework to carry all of the orchestration.[2]
The important idea for me is not that a filesystem is a new universal runtime. It is that structure can carry meaning. The right file in the right place can make the work inspectable, portable and understandable to both a person and a machine.
There is also a hard boundary. A folder can describe permissions, record an approval and carry evidence of a control. It cannot, by itself, establish identity, enforce access, prove compliance or suspend a worker. Files can carry the work contract and the governance record. They are not a substitute for the systems that enforce them.
One prompt. One folder. One harness.

One prompt, one folder, one harness
There is a reason the simple front door is attractive.
A team member should not need to understand the entire orchestration model to start a piece of work. They should be able to attach an approved source pack, point the harness to the entry file and state the outcome they need.
Behind that front door, the harness may still:
- find and validate the inputs;
- load a reusable skill;
- call tools and deterministic scripts;
- move through declared workflow stages;
- ask for independent challenge where it adds value;
- retain evidence and decision records; and
- stop for human judgement at the right gate.
One interface does not mean one model call. It does not even mean that one reasoning role performed everything. It means the complexity has been composed around the work rather than handed to the user.
The model is only the engine. Most of the design work sits around it: roles and responsibilities, inputs, authority, hand-offs, ownership and acceptance. If those are weak, adding another agent usually adds another place for the system to fail.
This is consistent with the distinction Anthropic makes between workflows, where models and tools follow predefined code paths, and agents, where the model dynamically directs its own process. Its guidance is to start with the simplest solution and add complexity when the outcome justifies the trade in latency and cost.[3]
The same principle appears in the way reusable skills are now packaged. Anthropic describes Agent Skills as modular, filesystem-based capabilities that bundle instructions, metadata and optional resources, and that can be composed when the task requires them.[4] A specialist capability does not always need a specialist worker waiting around it.
Agent, skill, tool, workflow or framework?
The vocabulary becomes confusing because these components are often presented as competing platforms or stages of maturity. I find it more useful to ask what each component is responsible for.
A rule is appropriate when the input and transformation are stable. A tool provides a defined interface to information or action. A skill packages reusable bounded competence. A workflow declares steps, gates and hand-offs. A framework provides shared concepts, roles and artefact contracts across a family of work. A harness gives the model an execution loop, context, tools and permissions.
A persistent worker becomes useful when the software needs ongoing purpose, context, judgement, authority or performance history. A team becomes useful when distinct expertise, concurrency, independence, access or operational ownership creates enough benefit to pay for coordination.
This is not a ladder. Each is a design choice.
Choose the smallest responsible component

Consider a wholly fictional example. Harbourlight Mobility receives a 60-page public consultation paper and wants a five-page impact brief for a leadership discussion.
Harbourlight Mobility · One bounded impact brief

At that point, creating a named research worker, analysis worker, writing worker and review worker may add more theatre than value.
Now change the conditions. The monitoring becomes continuous. The software can notify business owners or act in other systems. The performer and assurer need different access. Work must meet a service level. Performance and cost need to be managed over time.
The capability has not simply become more sophisticated. The responsibility has changed. Persistent, governed workers may now be the smallest responsible composition.
Current OpenAI guidance offers multi-agent coordination for complex work that divides cleanly into independent workstreams.[5] That is exactly the point: it is useful when the work earns it, not because worker count is a maturity badge.
Repeatability without rigidity
The obvious objection to a framework is that it can standardise away the thinking we wanted AI to contribute.
I do not want every answer to be the same. I want every answer to be traceable to the same discipline.
Keep the outcome, source boundary, mandatory gates, evidence, acceptance criteria and stop conditions constant. Keep the reasoning path, questions, approved sources, tools, models, creative alternatives and non-mandatory response structure adaptable.
The framework should keep the discipline constant without deciding the answer in advance.
That distinction matters. The purpose is to remove avoidable setup, ambiguity and control work so that more attention can go into judgement and creation. It is not to automate thinking into one approved shape.
I also need to be honest about the evidence. I have deliberately worked ahead on parts of the operating model and governance. Some of that thinking is now structurally tested; I have not yet proven every control through observed behaviour and sustained operation.
I have seen this discipline make bounded internal work clearer and more inspectable after evidence, challenge and human review. I have not yet established that an ordinary teammate can pick up the method cold and produce repeatably accepted outcomes without me. Until that test exists, this remains a practitioner's correction and a design hypothesis—not a productivity claim.
I have also been testing this through Bob, my supervised AI collaborator. We have added specialist roles, approval routes and review lanes. At times that made the work clearer. At others Bob was still the hidden integrator, or the documentation ran ahead of what we could prove in operation. I have examples of work being held when evidence or authority was missing. What I do not yet have is a clean record of repeatably accepted outcomes or an earned increase in autonomy. That distinction matters.
When does work become workforce?
There is no universal number of steps, tools or model calls that turns work into a workforce. I use a set of observable questions instead:
- Does the software act beyond preparing a human decision?
- Does it hold differentiated identity, authority or permissions?
- Does responsibility persist across sessions, cases or human owners?
- Must the work run continuously, concurrently or to a service level?
- Could failure create material, regulated, customer or production impact?
- Must challenge or assurance be independent from the performer?
- Must performance, cost, capacity, trust or improvement be managed over time?
- Does another worker need a distinct context or segregation boundary?
This is a proportionality test, not a compliance score. One weak “yes” does not automatically require a digital workforce. One strong authority, consequence or segregation boundary may justify worker governance immediately.

Does the software now hold sustained authority, consequence or responsibility?
On the bounded side of the threshold, govern the work: outcome, sources, method, evidence and human acceptance.
Across the threshold, govern the worker as well: identity, permissions, delegation, trust, assurance, monitoring, intervention, performance and lifecycle.
This is why I am not arguing against agent teams. In sustained digital work, the operating model is not optional decoration. It is the difference between a clever demonstration and a responsibility the enterprise can actually own.
The management problem begins
The mistake is not building agents. The mistake is building a workforce before we have defined the work, or pretending that a tidy folder is enough after the software begins to hold responsibility.
Start with the work contract. Add the smallest component the outcome earns. Keep the human-facing interface as simple as it can be. Preserve evidence and acceptance. Then pay the cost of persistent identity, separation and workforce governance when the authority, consequence or operating conditions justify it.
Govern the work before you build the workforce. When the work becomes sustained, delegated or consequential, govern the worker as well.
Choosing the simplest responsible way to perform the work is only the first decision. Once software begins to hold sustained responsibility, the model is still only the engine. The harder question is no longer how many agents we can build. It is how roles, authority, inputs, evidence, human control, lifecycle and ownership work together around them.
That is where this article hands over to Beyond Agents: The Digital Workforce Operating System. This piece starts with the smallest responsible composition. The wider conversation is about what it takes to manage the workforce once that responsibility persists.
Sources
- NVIDIA, “What Is an AI Factory?”, accessed 23 August 2026.
- Jake Van Clief and David McDermott, “Interpretable Context Methodology: Folder Structure as Agentic Architecture”, arXiv:2603.16021v2, 18 March 2026.
- Anthropic, “Building effective agents”, 19 December 2024. Used only for the workflow-agent distinction and proportional-complexity guidance.
- Anthropic, “Agent Skills”, accessed 23 August 2026.
- OpenAI, “Model guidance”, accessed 23 August 2026. The referenced multi-agent feature is described as beta.
What do you think?
If this sparked an idea—or you see it differently—I’d like to hear it.
Send me a note