Benjamin Johnson
All writing

Beyond Agents: Why Your AI Strategy Needs an Operating System

The model is not the hardest part. The hard part begins when agentic AI takes on work your organisation must stand behind.

Benjamin Johnson13 min read

Written by Benjamin Johnson, with AI-assisted research and editorial support. Illustrations generated with AI.

The demo is over.

Who owns Monday?

A person pulls back an oversized AI demonstration interface to reveal the human-led organisation behind it: shared work, review, access boundaries and recovery responsibilities.
The demonstrationShows what the capability can do.
The organisationOwns how the work is run.
The consequenceSomeone must stand behind the result.
Getting it to work is one thing. Running it every day brings a few more questions.

My journey with AI began well before ChatGPT. It started with machine learning: exploring image recognition, facial expressions, sentiment analysis and object detection. I was interested in what these systems could recognise, interpret and help us do differently.

Then came chatbots, followed by tools such as ChatGPT and Stable Diffusion. The possibilities expanded into conversation, writing and image generation. Alongside that, my thinking about services evolved—from intelligent managed services to AIOps, and now into agentic operations, agents and the systems around them.

The way I use AI has changed too. Simple research questions and requests for summaries have opened into longer pieces of work: developing a product or service offering, preparing an RFP response, or assembling market intelligence. Increasingly, I am exploring workflows in which agents research, analyse, draft and challenge each other’s contributions, with people setting the brief and judging the result. That brings questions about sources, access, hand-offs and ownership into the work itself.

In Are We Overcomplicating Enterprise Agentic AI?, I explored where to begin: define the outcome, then choose the simplest combination of tools, skills, workflows and agents that can responsibly deliver it. A bounded task may need very little machinery. Persistent workers and teams need to earn their place through the responsibility they carry.

This article picks up where that one left off. When agentic AI is expected to keep working, exercise delegated authority or carry consequences across cases and hand-offs, we have an ongoing management responsibility. How do we operate it day after day? Who owns the result? What happens when it is wrong?

That brings me to the question at the centre of Beyond Agents:

What changes when agentic AI becomes part of the team?

Here I mean sustained responsibility for work with an owner, a customer, a cost, a consequence and sometimes a regulator waiting on the other side. A single task does not automatically require a digital workforce; the authority, consequences and operating conditions determine how much governance it needs.

My view is that agents carrying that responsibility should be treated as an extension of the team. We should bring the same seriousness to their role, induction, access, supervision and performance that we would bring to a human colleague. The controls will differ: a model can change between versions, access can be revoked instantly, and a convincing answer can still be wrong. People remain accountable for the work entrusted to the system.

That is the gap I want this series to examine: whether we are giving as much attention to managing digital labour as we give to building the agents that perform it.

The Monday after the demonstration

In a demonstration, we can select the data, narrow the task and keep the people who built the system close at hand. A proof of concept can answer a useful question about capability while leaving much of the operating model untested.

Getting into production means fitting that capability into the organisation: identity and access, data boundaries, security review, change processes, service ownership, support and recovery. Keeping it useful after launch is the Day-2 challenge.

There is progress, alongside a gap in results. McKinsey’s August 2026 survey reports more organisations scaling AI, while the share reporting enterprise-level earnings impact remained broadly unchanged. Its high performers more often redesign workflows and measure impact. Those findings show an association; they do not establish why a particular pilot stalls.

My concern is that we too often leave the organisational questions until after the demo. I will admit to something of a love–hate relationship with security and cyber leadership over the years. When you want to move an idea forward, scrutiny can feel like friction. What has changed my approach is bringing security, governance and risk into the scope and business case from the beginning, with those leaders involved in shaping the proposal.

That makes the path clearer: what are we testing, which data and systems are in bounds, who can authorise an action, and what would have to be true before this could become a live service? It does not guarantee production approval. It means the POC can test more of what matters, instead of discovering the organisation’s requirements after the demonstration.

Demand changes. Systems become unavailable. A customer asks the question nobody prepared for. Two apparently sensible actions conflict. Costs rise quietly. A policy changes. An upstream source turns out to be wrong. The agent needs to hand work to another agent—or back to a person— and the receiving worker needs to understand what happened and why.

Better prompts can help with parts of this. The wider problem needs an operating model.

Someone still has to decide:

Management decisions

Someone still has to make the call.

Set the boundaries

  • who owns the outcome;
  • what the digital worker is allowed to do;

Control the action

  • which evidence must exist before it acts;
  • when a person must approve, intervene or take command;

Keep it accountable

  • how quality, cost and customer impact are measured;
  • how the worker learns without silently changing itself; and
  • when it should be restricted, retrained, suspended or retired.
People remain accountable across all three areas. A proposed change is not an approved change.

We already understand these questions when the worker is human or the service is conventional software. AI has not made them disappear. It has made the boundaries less visible.

An agent is not yet a workforce

I use the term digital worker deliberately.

An agentic system typically combines a model with an execution harness, instructions, context and tools. The harness manages the execution loop and tool use; the model can select actions in response to the task and results. Anthropic’s engineering guidance distinguishes that dynamic control from workflows with predefined paths. The model is one component of the working system.

I use digital worker for an agentic capability given a defined role in the organisation, with explicit authority, a human owner and a managed lifecycle. That is an operating-model definition, not a claim that an agent is a person.

That distinction matters because capability is not authority. A system may be technically able to complete a task and still lack permission, sufficient evidence or the right conditions to do it now.

A useful starting point is to ask what we would need to know before putting a new colleague to work:

Worker requirements

More than a name and a login.

  • Identity and ownership: Which worker and version is this, and which person is accountable for its use?
  • A job description: What outcomes does it support, which capabilities does it need, and what work falls outside its role?
  • System and tool access: What may it read, change or invoke, for which purpose and under whose authority?
  • Working rules: Which policies, evidence requirements, approval points and escalation routes must it follow?
  • Behaviour and communication: How should it explain uncertainty, collaborate and accept challenge? A recognisable personality can help people work with it; personality earns no permissions.
  • Performance and quality: What KPIs and acceptance criteria matter—accuracy, service quality, timeliness, cost, rework and appropriate escalation? Activity alone is not success.
  • A lifecycle: How is it inducted, evaluated and qualified for its role, reviewed over time, and requalified, restricted or retired when something changes?
Requirements around one worker. The organisation’s seven DWOS pillars follow in the next section.

We can borrow management discipline from human teams while designing controls for how agents actually behave. A role description states expectations; permissions, evaluations and operational checks must make those expectations real.

This is also the lens I am bringing to Bob and my personal agentic team. Giving agents names and specialisms is the easy part; defining their boundaries, reviewing their work and owning their mistakes is the management work. Bob’s team introduction offers a practical companion to these questions—not evidence that the enterprise model is already proven.

One useful agent is not a workforce. A workforce emerges when humans, digital workers and services operate together through explicit roles, hand-offs and decision rights.

The operating system around the work

My current working model is a Digital Workforce Operating System: a platform-neutral management and control layer for mixed human, agent and service work.

I do not mean an operating system in the narrow software sense, and I am not claiming to have invented the phrase. I mean the organisational machinery that allows digital labour to be governed, measured, trusted, improved and—when necessary—stopped.

The model has seven connected pillars:

  1. Governance — accountability, policy, decision rights, approval and audit.
  2. Economics — the fully loaded cost, capacity and value of producing an accepted outcome.
  3. Agent HR — how digital workers are selected, inducted, certified, reviewed, developed, suspended and retired.
  4. Mission Control — how work is allocated, observed, escalated and recovered during live operations.
  5. Hybrid Operating Model — how people, digital workers and services work as one accountable organisation.
  6. Continuous Organisational Learning — how observations become evaluated improvements, certified changes and shared knowledge.
  7. Confidence, Trust & Assurance — how we decide whether this worker should do this task now, whether it has earned reliance over time, and whether we can prove the decision was governed.

The value is not in treating these as seven independent workstreams. It is in the way they constrain and strengthen one another.

Governance defines the rules. Confidence, Trust & Assurance tests a proposed delegation against those rules and the current context. Mission Control applies the decision and watches what happens. Agent HR and organisational learning update the worker’s evidence. Economics and the hybrid operating model determine whether the arrangement is viable and who carries the remaining responsibility.

If one pillar is missing, the weakness appears somewhere else—often after the demonstration.

Confidence is not trust, and trust is not assurance

These words are often blended into a general claim that an AI system is “trusted”. I think that is too vague to operate.

For the model I am developing:

  • Confidence asks: should this worker perform this task now, under these conditions?
  • Trust asks: has this worker earned revisable reliance over time and within this scope?
  • Assurance asks: can we show the evidence, controls and decisions that make that reliance defensible?

A high score should never override identity, authority, certification, prohibited states or a required human decision. Nor should a worker automatically inherit trust because another trusted worker selected it.

The numerical models and autonomy thresholds behind these questions still require calibration, challenge and controlled testing. The operating principle comes first: trust must remain scoped, evidenced and revocable.

The human workforce is part of the architecture

The hybrid workforce cannot be designed as a digital system with people left around the edges to approve exceptions.

The latest releases make this question more immediate. OpenAI describes improvements in computer and browser use with GPT-6 Astra, while Anthropic describes stronger coding and long-running knowledge work with Claude Fable 5.1. These are vendor capability claims, not proof that an organisation can safely delegate the same work. We do not need to settle the AGI debate to ask whether we are handing over more authority than we can meaningfully supervise.

Human in the loop has to mean more than a person clicking approve. That person needs the expertise, evidence, time and authority to challenge a decision, stop execution and take over. Some actions should wait for explicit approval; bounded, lower-risk work may run under monitoring. A human name on an escalation chart is not a control if nobody can respond in time.

There is no useful universal ratio of humans to agents. I would size the human team around the consequences of error, review effort, peak exception demand and recovery workload—not the number of agents it can nominally watch. Ten quiet agents and ten agents escalating at once are very different staffing problems.

If digital workers absorb routine execution, we need to decide how people will still build judgement. If recovered capacity is treated only as a headcount target, we may make today’s work cheaper while quietly weakening tomorrow’s expertise.

Can the business still operate when its digital workforce cannot?

The AI stops. The work doesn’t.

An unplugged connector leaves digital capability dark while work continues to arrive. Two people use an independent procedure to handle priority cases; other work waits and an unauthorised request stays behind a closed boundary.
ContinueCritical incident and urgent access.
WaitRefund, routine question and summary.
Still stoppedUnauthorised privilege change.
Less capacity. The same responsibility.
What this example is showing

In this fictional six-request example, each of two qualified people can actively handle one case. One continues a critical incident. The other defers a refund review to take an urgent access request within their existing authority. The refund, routine question and weekly summary wait; an unauthorised privilege change stays stopped.

No case is shown as completed. This assumes independent access, current procedures and retained skills. It illustrates prioritisation, not a staffing ratio, service promise or production result.

If model access is withdrawn, a provider goes down or the execution harness becomes unavailable, switching models may not be enough. Several agents can share the same point of failure. The continuity plan needs usable procedures and business knowledge outside that dependency, people who retain the skills and system access to act, and a rehearsed way to stop automation and reconcile unfinished work.

That does not mean a smaller human team must reproduce every automated task at the same speed. It means deciding which critical services must continue, what reduced service is acceptable, what can wait and for how long—and testing whether the retained team can actually deliver it. A procedure nobody has practised is not a credible fallback.

The right comparison is also not salary versus token cost. A digital worker has a fully loaded cost: models, platforms, integration, observability, evaluation, security, governance, human supervision, exception handling, recovery and ongoing change. A human role has its own fully loaded cost and contributes judgement, relationships, learning and resilience that do not fit neatly into a transaction count.

The useful question is whether the combined workforce produces better accepted outcomes—at an understood cost and risk—while preserving the capabilities the organisation will need next.

That is a management decision, not merely an automation calculation.

Why I am publishing this now

These questions also run through my professional work on enterprise services and operating models. As I explore how AI fits into service delivery, I keep coming back to what an organisation would need to take responsibility for it: clear ownership, dependable service, security, an understood cost and people who can step in when circumstances change.

That enterprise lens shapes the questions I bring to this work. How does a digital worker fit across existing teams, suppliers and business functions? Who owns an outcome that crosses those boundaries? What changes for the people delivering the service, and how do we preserve their judgement as more execution moves into software? Technical support is one place to examine those questions. Customer service, finance, marketing and product development need answers too.

Alongside that professional perspective, I have been building and operating Bob, my personal AI assistant, on my own infrastructure. Bob gives me a practical place to explore roles, delegation, review and recovery. It also makes the gaps hard to ignore. What I learn there helps me sharpen the questions; it does not establish that the same approach is ready for enterprise use. This series is my personal perspective, and the examples I share here come from that personal work and public sources.

Bob shares his own introduction to the team in Meet the Team: What My AI Colleagues Taught Me About the Enterprise Operating Model. It offers a companion perspective on why the roles exist, where hand-offs become difficult and how easily coordination can be mistaken for delivery. Those are exactly the questions I want this operating model to help us examine.

The model is rarely the hardest part.

The hard part is knowing what the system is allowed to do, what it actually did, whether the result was acceptable, what it cost, who can challenge it and how to recover when it is wrong.

Where this goes next

This article opens a wider series. The forthcoming paper, Beyond Agents: The Digital Workforce Operating System, will bring the seven pillars together. Supporting articles will take the practical questions one at a time:

  • Governance, confidence and trust: When should a worker act, what changes when it delegates to another agent, and what evidence makes the decision defensible?
  • Agent HR and the hybrid organisation: How do we introduce, review and develop digital workers—and how do people build expertise when AI takes on more of the junior work?
  • The real economics: What does an accepted outcome cost once models, supervision, security, coordination and rework are included? When could that support an outcome-priced service?
  • Day 2 and Mission Control: Who watches the work, handles exceptions and takes responsibility when a hand-off fails—and can critical services continue without the model or harness?
  • Learning and sustainability: How do we improve the system without allowing uncontrolled change, and sustain its human, economic, operational and environmental foundations?

Across the series, I will also explore how these responsibilities change between a large enterprise and a small business, where one person may wear several of the management hats.

These are not finished standards or claims of conformity. They are practitioner propositions being made explicit so they can be tested, challenged and improved in public.

The shift I am interested in is not from people to agents.

It is from isolated agents to governed digital labour—and from two separate workforces to one accountable organisation.

If agentic AI is becoming part of the workforce, what operating system does your organisation need around it?

Read the companion

Meet the Team: What My AI Colleagues Taught Me About the Enterprise Operating Model