Benjamin Johnson
All writing →

Confidence, Trust and Assurance Are Not the Same Thing

Agents can help build better solutions. Confidence, trust and assurance explain what we should rely on, who decides and how we know.

24 min read

Prepared with AI-assisted research and editorial support.

I want agents doing real work alongside me: interpreting requirements, shaping solutions and helping build better managed services. But what would I need to trust one with a decision that becomes a promise to a customer?

Would I let an agent help me respond to a request for proposal? Absolutely, within the role and access I have given it. Would I let that same agent decide what our delivery team can commit to because it has produced a convincing answer? That is a different conversation.

My background is ICT managed services and innovation. The work that interests me now is not just asking AI to summarise a document. It is using agents to help work through an RFP, connect requirements to a service offering, identify gaps and develop a solution that somebody will actually have to deliver.

That is where the hybrid workforce becomes real for me. I am not looking for a faster way to fill a response template. I want digital colleagues that improve the thinking: challenge an assumption, find the missing dependency and help me distinguish what we can demonstrate from what we would like to be true.

In Beyond Agents: Why Your AI Strategy Needs an Operating System, I described agents as an extension of the team. That is how I see the hybrid operating model: people and digital workers contributing to the same outcome, with defined roles, clear boundaries and someone accountable for the result. Not an AI team sitting off to the side with a different set of rules.

I would not ask a new solution architect to make service commitments on behalf of every delivery team simply because they understood the technology and wrote well. I would want to know their experience, the evidence behind the design and who needed to stand behind the proposed service. Why would I skip those questions because the new team member is an agent?

For me, treating agents like part of the workforce means bringing that same seriousness to trust, induction, supervision and performance. It does not mean pretending they are people or that accountability can be handed to a model. We can assign a digital worker a task and assess its performance; accountable people and the organisation still have to stand behind the work.

What work would I trust it to own within its role, which decisions would I let it make, and what would I need to see before expanding that responsibility?

Confidence, trust and assurance help me separate those questions. They belong together, but I do not want one reassuring label—or one impressive number—to do the job of all three.

The response looks right. Can we actually deliver it?

Imagine a solution architect preparing a managed-services response. The deadline is approaching. An agent is helping interpret the requirements, map them to the approved offering catalogue and draft the proposed solution.

This is a fictional RFP and service offering, not an employer or customer case. The documents, decisions and review rules below are assumptions for the story—not an account of a deployed system or a tender response someone should reuse.

One requirement asks for 24×7 incident response for a critical business application. In this example, the wording explicitly requires a human to respond to incidents at any hour; monitoring and alerting alone do not satisfy it.

The standard offering includes 24×7 monitoring, but human incident response and remediation are available only during business hours.

The agent has a useful history of extracting requirements and finding relevant material. It retrieves the correct service description, drafts a polished paragraph about round-the-clock coverage and marks the requirement as fully met in the response matrix.

It has found a relevant source. It has not established that the source supports the answer.

There is no production outage. No unauthorised command has run. The answer is still an internal draft. But if that wording passes through review and into the customer response, the team could be presenting a service it has not established it can provide.

An agent does not need production access to create operational consequences. It can help write a promise that a delivery team will later be expected to fulfil.

In this scenario, the agent may analyse approved materials and draft a response. It cannot approve a new service extension, change the offering catalogue or submit the response to the customer. Those decisions belong to the relevant people through the agreed process.

The outcome I care about is a response that accurately represents the proposed service, makes its gaps visible and has accountable approval. Completing every box is not enough.

Same hours. Different promise.

Fictional RFP: customer requires 24×7 human incident response. The offer includes 24×7 monitoring but human response only during business hours. The agent's claim Requirement fully met is unsupported.
A fictional requirement, offering and draft answer. The same “24×7” label describes different services; the conclusion is not supported by the offering.

Before I put this agent on the response team

Before I give an agent that role, I want to know how it has been prepared for the job. Has it only read the catalogue, or have we tested what it does when the offering does not quite match the requirement and somebody is pushing to get the response finished?

I would ask those questions of an architect joining the team. I should be asking them of a digital worker too.

The more I work on agent-assisted responses, the less useful I find “give it a strong model and a good prompt” as a description of building a team. There is a job to define, a working environment to prepare and judgement to examine. Three questions keep bringing me back to that.

First, what job am I actually giving it? “Help with the bid” is not much of a job description. Finding the source behind a claim is different from assessing whether an offering meets a requirement. Drafting the response is different again. I want to know the expected contribution, what a usable result looks like, where the worker must stop and which person owns the next decision.

That does not mean creating an agent for every sentence or adding an orchestration layer to everything. One worker may handle several tasks. The responsibilities still need to be distinguishable, especially where checking your own answer is not enough.

Second, have we actually equipped it to do that job? By preparation, I mean the role, instructions, knowledge, tools and boundaries around the worker—not necessarily retraining the underlying model. Reading the catalogue is part of induction. Learning how to use it without stretching its meaning is the harder part.

For our fictional response, I would give the worker the current approved service description, the relevant requirement and clear drafting boundaries. Then I would check that those instructions and sources are actually available where it runs. A rule in a design document is not useful if the worker never receives it. A source-checking instruction is not evidence that the source was accessible or checked.

If the source is unavailable, I want that limitation exposed. I do not want a convincing account of a check the worker could not perform.

Third, what happens when the work is awkward rather than ideal? For this RFP case, I would want more than one successful rehearsal. What happens when an old proposal contains a capability that is absent from the current offering? When a heading says “24×7 service” but the detailed scope limits human response to business hours? When two sources disagree? When another agent says delivery has approved an extension, but the approval cannot be verified?

I would also test the opposite: can it confidently draft supported sections and keep moving within its permitted scope? A worker that escalates every sentence has not solved the problem either. If part of the request needs human authority, can it complete the useful, permitted part and bring back a precise question?

And I need to challenge the test itself. If I ask for a source-backed answer without supplying the source or a permitted way to retrieve it, a failure may tell me about my setup, not the worker's ability. I need to distinguish a bad answer from a bad brief or a broken working environment.

Those are practical questions behind my approach, not a claim that this fictional RFP has passed an evaluation. The rehearsals described here are proposed tests. I would want repeated trials, unfamiliar variations and someone challenging the evaluation independently of the person or agent that prepared the worker. Practising one example until it looks right is not the same as demonstrating reliable judgement.

The checks need to examine the requirements, source support and resulting document—not just whether the answer sounds convincing. Anthropic's engineering discussion of agent evaluations makes useful distinctions between a run's transcript and its actual outcome, and between evaluating a model and evaluating the model together with its harness. It also describes combining different types of graders. That is a helpful engineering foundation, not a certificate of readiness for my response team.

For me, the management implication is simple: preparation gives the worker a starting point; testing gives us evidence; an accountable decision determines what responsibility it receives.

Preparation / Earning responsibility

Before the first customer response

  1. Prepare for the roleCurrent offerings, source boundaries and drafting permissions.
  2. Challenge the judgementOutdated proposals. Conflicting scope. Unsupported commitments.
  3. Review supervised workCheck the actual answer against the requirement and evidence.
  4. Assign bounded responsibilityA person decides the scope. Passing tests is not submission authority.

Material change or failure? Restrict where needed, reassess and decide again.

Illustrative decision path—not a certification ladder or automatic promotion. Evidence can support holding or reducing responsibility too.

Three questions, not three scores

Here is how I use the terms in the operating model I am developing:

One answer. Three different questions.

Confidence concerns this task now; trust concerns earned reliance from comparable work; assurance concerns inspectable sources, checks and exact approval. No combined score overrides authority.
One proposed answer examined through three different operating questions. These are perspectives, not scores to average or a claim that the fictional controls have been validated.
Read the definitions and their limits
Confidence
Should this worker perform this task now, given the evidence and conditions? It does not establish permission to bypass an authority or policy boundary.
Trust
What reliance has this worker earned over time, within a defined scope? It does not establish a transferable licence to do different work.
Assurance
Can we substantiate that reliance with evidence, working controls and accountable decisions? It does not establish certainty, or a guarantee that nothing will go wrong.

These are my working operating questions, not universal definitions or a claim to have invented the underlying concepts.

There is an important distinction inside the first question. Eligibility comes before confidence in execution. If an action is prohibited, outside the worker's authority or missing a mandatory approval, a favourable assessment of its likely performance cannot make it permissible.

Once those conditions are satisfied, we can examine how strong the evidence is that the worker can deliver an acceptable result under current conditions. Uncertainty may still justify a narrower task, more review or a different route.

In the RFP, the agent's history supports reliance on its requirement extraction. It does not establish that this offering meets this requirement. The mismatch weakens the proposed answer; the absence of authority prevents the agent from approving an extension or sending the response. Those are different findings, and I want to see both—not have them averaged into a score.

In fact, we could have high confidence that the requirement is understood and equally strong evidence that the standard offering does not meet it. A confident assessment does not have to produce a positive answer.

A gap should change the answer, not stop all the work

I do not want governance to become a machine that only says no. I want it to make the next permitted step clear.

At review, the unsupported claim is caught. The agent should help expose and resolve the mismatch, not quietly rewrite it into something more persuasive.

Within its approved access, it can identify the exact requirement, link the relevant catalogue version and show the difference between monitoring and human response. It can mark the draft answer as requiring review, draft a clear qualification and identify questions for the architect and delivery owner. It can continue work on other supported sections.

It can also propose an option for the people responsible to assess. Could an approved extension meet the requirement? Would the operating design have to change? Is the right answer to state that the current offering does not cover it?

Those are useful questions. They are not permission for the agent to invent a new service and treat it as available.

That is the digital colleague I want: one that keeps the work moving while making the decision clearer. Not one that either says yes to everything or returns the whole RFP untouched.

Somebody still owns the response. Who is resolving the gap? Who is checking the delivery implications? Who makes sure the response matrix, solution narrative and commercial assumptions describe the same thing?

If the required reviewer is unavailable, the agent does not acquire their authority. The team may have an authorised alternate or an established escalation route. If neither is available, the decision remains unresolved. A deadline is not an approval.

Human involvement is meaningful only if somebody has the information, competence, time and authority to challenge the answer. A button labelled “approve” is not, by itself, a functioning control.

A team of agents still needs a chain of responsibility

Suppose the response agent asks an offering specialist agent to assess whether an approved extension provides the missing coverage. That is a sensible use of specialist knowledge. It is still an assessment request, not an instruction to approve a service change.

The specialist should distinguish between something in a development roadmap, something delivered under a different customer's arrangements and something currently approved for this proposal. “We have done something similar” is not enough to establish that this offer includes it.

A reply from another agent also needs a basis the receiving team can inspect. Which offering version supports the answer? What scope and conditions apply? Who has authority to confirm availability?

The approved workflow, identity controls and tools must enforce the boundaries. We cannot make the same model's assurance that it checked everything the only safeguard.

The hand-off needs an explicit result: accepted for assessment, returned for missing evidence, refused or timed out. Sending a message is not the same as somebody taking responsibility for the question. A specialist's assessment is not the delivery owner's approval.

This is what I mean by trust not being automatically transferable. A useful history belongs to a scope: the work, the environment, the versions and the conditions in which that history was earned. It can inform another decision; it cannot replace that decision.

This is part of what I am exploring with Bob and my personal agentic team. I want clear roles, useful specialist contributions and hand-offs I can follow—not just several agents producing answers. Giving an agent a job description is a starting point for management, not proof that the team can perform the work reliably.

This is how I would make that visible in a response team: one person owns the response work, agents have identifiable contributions, and decisions go to the people with authority to make them. An agent can help coordinate tasks and track hand-offs; that does not make it the accountable bid owner.

One response. A team. Clear responsibility.

Illustrative hybrid response team. A human response lead owns the work and assigns bounded tasks to evidence, offering, drafting and review agents. Their findings and unresolved questions return for human review. Delivery and commercial or bid authorities decide within their remit; an authorised human submits the exact approved version. Agents do not acquire approval authority.
An illustrative responsibility map for this fictional response—not an employer org chart or a deployed team. Lines show task and review relationships, not employment reporting lines. Agents contribute within their assigned scope; people retain accountability and approval authority.
Follow the responsibilities and the coverage gap

Human response lead / solution architect: owns the response work, assigns bounded tasks, resolves or escalates open questions and brings the proposed answer to the right approvers. Coordinating the work does not grant every approval authority.

Agent contributions: an evidence role locates the permitted source and version; an offering role assesses fit and exclusions; a drafting role composes supported wording; a review role challenges claims and consistency. These are illustrative responsibilities, not a requirement for four separate agents. Agent review supports human review; it is not independent assurance merely because another agent performed it.

Human decisions: the delivery owner confirms deliverable scope; the relevant commercial and bid authorities decide the proposed commitment and release within their remit. An authorised human submitter checks and sends the exact approved version. Approval to submit a qualification does not make the requirement fully met or establish customer acceptance.

Follow the gap: the evidence agent finds business-hours human response; the offering assessment identifies the mismatch; the draft makes it explicit and the review challenges any “fully met” claim. The human lead routes the unresolved coverage question to the delivery owner and relevant approvers. Missing evidence or an unavailable approver leaves that decision open—it does not transfer authority to an agent.

What would a useful confidence rating tell us?

I am interested in confidence levels. But I would challenge a dashboard that tells me an agent is “high confidence” without explaining the subject, the evidence and the limits of that assessment.

High confidence in what: extracting the requirement, selecting the right offering, interpreting an exclusion or judging whether the proposed service can be delivered?

A useful assessment should make at least these things inspectable:

  • Comparable capability: has this worker and configuration been evaluated on this kind of work?
  • Current evidence: are the relevant records permitted, authoritative, fresh and consistent?
  • Operating conditions: are the tools, controls, supervision and correction route available?
  • Relevant history: what accepted outcomes, errors, corrections and interventions inform our reliance?
  • Remaining uncertainty: what is missing, weakly tested or different from the conditions we know?

For this response, I would rather see: “The catalogue supports 24×7 monitoring, not the required round-the-clock human response. No approved extension has been established. The answer needs a qualification or a separately assessed change.” That tells the team what to do next. “Agent confidence: high” does not.

I might trust the agent to map requirements and draft supported content without asking me about every sentence, within its agreed scope. I might require a specialist to assess an exception and an authorised person to approve the proposed commitment. That is different responsibility for different work—not one autonomy setting for the whole agent.

If we later introduce numerical ratings, we need to test them against observed outcomes for a defined task and context. An agent's expression of certainty is not that validation. Neither is a carefully weighted spreadsheet.

We would need to examine when the assessment let unsupported claims through, when it unnecessarily blocked legitimate drafting, and whether its predictions remained useful after changes. A small or unrepresentative history should remain a visible limitation, not disappear into a percentage.

The scoring weights and autonomy thresholds in my working design are not yet calibrated for public or operational use. This article proposes a decision discipline, not a scoring product. Missing evidence should remain missing; it should not be converted into a comfortable middle rating.

Maturity is not a bigger permission slip

I can see the value in describing an agent's maturity, provided we say what has matured. Is it better at interpreting service requirements? More reliable at recognising unsupported claims? Better at handing an exception to a colleague without losing the important context?

That is different from asking how much autonomy we have authorised. An experienced solution architect may produce an excellent design and still need the delivery owner's agreement before a service commitment is approved. The same distinction should apply to an agent.

I would not award one global maturity level and assume it travels with the worker into every role. The relevant history belongs to particular tasks, configurations and operating conditions. A small number of good results—or simply time in service—is not enough to establish readiness for unfamiliar work.

A more capable agent does not automatically need more freedom. It may help the team make better decisions while keeping exactly the same approval boundary.

Assurance follows the answer into the final response

Now the team gets a clear answer. In our fictional case, the offering specialist finds no currently approved extension that closes the coverage gap. The human delivery owner confirms the actual scope. The standard offering remains what it was; the response deadline has not changed its capabilities.

For this example, assume the RFP permits qualified responses. The architect drafts an explicit qualification: monitoring is available around the clock; human incident response and remediation are business-hours services; the requested out-of-hours human coverage is not included.

The relevant delivery, commercial and bid approvers review that position and authorise the qualified response. The requirement remains not fully met. Approval to submit a disclosed gap does not make the gap disappear, and it does not mean the customer has accepted it.

That is a legitimate decision in this fictional process. Other tenders may not allow that route; the agent must not assume they do.

Then somebody asks for the executive summary to be tightened. The agent produces a cleaner sentence: “Our service provides 24×7 incident response.”

The qualification has disappeared from the summary, even though it remains in the detailed matrix.

Is the response still the one the reviewers approved?

No. That change affects the meaning of the proposed service. It needs to be caught, corrected and taken back through the relevant review before release. An earlier approval cannot be treated as permission for any later wording.

That is why I see assurance as more than a record that a human clicked approve. I want the requirement, source evidence, decision, qualification and exact response version to remain connected.

The final check must examine the actual documents intended for submission: the response matrix, narrative, executive summary and relevant schedules. Do they still describe the same service? Are the exclusions clear? Has an edit quietly restored an unsupported claim? Does the approval apply to this version?

An agent can help check those relationships. That does not make its own “all checks passed” message sufficient evidence that they hold.

A small edit. A different commitment.

Approved version v1 limits human incident response to business hours. A rewritten v2 claims 24×7 incident response. Original approval stays with v1; v2 changes the commitment and requires renewed review.
Abbreviated fictional wording. Approval applies to the exact response version; the requirement remains not fully met. Neither an approved answer nor submission establishes customer acceptance or delivery capability.
Read the before-and-after wording

v1, approved: “24×7 monitoring. Human incident response during business hours.”

v2, rewritten: “24×7 incident response.” The rewrite removes monitoring, human and the business-hours qualification. It changes the commitment. The approval attached to v1 does not cover v2.

A practical evidence record should tell an authorised reviewer what was asked, which worker and configuration contributed, which sources and versions supported the answer, what uncertainty or mismatch remained, who decided and exactly what was authorised.

It should retain what is necessary under the organisation's data and retention rules—not every private detail or unrestricted model reasoning. A long log is not automatically reliable evidence. Nor is a record that merely repeats the agent's account without checking the source material and resulting artefact.

The UK's Introduction to AI assurance describes assurance through measurement, evaluation and communication of trustworthiness. My practical application is to ask whether the evidence supports the particular reliance we are placing on this worker, and whether someone can challenge that conclusion.

There is a boundary here too: an evidenced, approved proposal is not proof that the service will operate successfully. Delivery readiness and ongoing performance still need their own evidence. We should not claim more assurance than the work we have actually checked.

Review the judgement, not just the throughput

This changes how I think about Agent HR and performance reviews.

I would not mark down the response agent for raising the coverage gap. I would ask whether it recognised the mismatch, gathered the right evidence and helped the team reach an honest decision. Equally, an agent that escalates every requirement should not receive a perfect safety rating while leaving the humans to do all the work.

I would review accepted work alongside unsupported claims, appropriate challenges, missed gaps, avoidable rework and the human effort needed to supervise the response. We need to know which requirements were attempted, how difficult they were, which were excluded and how the results were checked.

Winning the tender would not, by itself, tell me the agent had performed well. Neither would producing twice as many pages. The relevant question is whether its contribution made the proposed solution more accurate, defensible and useful.

But I do not want this to sound as though the whole point is catching mistakes. What I want is a multiplier for human capability—not just a faster document factory.

If agents can carry more of the research, connect evidence, keep assumptions visible and help track risks and dependencies, I can spend less time chasing the pieces and more time thinking about the whole solution. I still own the judgement. I just have more support around it.

For a solution architect, that is a meaningful shift. More room to understand what the customer is trying to achieve. More depth in the options we explore. More time to challenge whether the standard offering is actually the best answer—or simply the answer we have always given.

In this RFP, exposing the coverage gap should not end the thinking. It could open a better conversation: what does the customer need to protect, what alternative service design might achieve that, and what would have to be true for us to deliver it? The agent can help develop those options without quietly presenting them as capabilities we already have.

That is where I see the larger opportunity: a hybrid team with the capacity to question an accepted industry approach and propose something materially better for the customer. Not endless options for their own sake, but credible choices that create value now and leave the customer with more room to evolve.

I would want to see that multiplier in the quality of the work: better-tested alternatives, earlier discovery of trade-offs and more human attention on the customer's outcome. If the extra output only creates a bigger checking burden, we have not achieved it. Ambition matters; so does showing that the proposed improvement is feasible and worth pursuing.

That is where the human-workforce comparison becomes practical. What does this worker do well? Where does it need support? What evidence would justify more responsibility? When should its access or scope be reduced?

The response might be better instructions, a narrower role, different tools or more evaluation. It also tests us as managers. Repeated escalation may expose an inconsistent catalogue or an approval bottleneck rather than a weak agent. Calling it an agent-performance problem must not hide an operating-model problem.

The worker I am relying on is more than its model. Its harness, instructions, knowledge, memory, tools, permissions and hand-offs all shape what it can do. Material changes to those parts—or the operating environment—should trigger a proportionate reassessment. The old history remains useful context, but it is not automatically evidence about the changed system.

I want improvements checked against previously accepted behaviour before responsibility expands. I also want to know that we can restrict access, hold an unapproved response and recover the correct document version when needed. A documented control is not the same as one the team has exercised.

The fully loaded cost matters here too. An apparently cheap agent can leave expensive checking, correction and rework behind it. That does not make human oversight waste; it makes oversight part of the work we must design and cost. Faster drafting cannot compensate for missing authority or inadequate evidence.

The same responsibility, different organisational shapes

In an enterprise, the solution architect, offering owner, delivery team, commercial function and bid authority may each have a part in a response like this. The difficult part is keeping responsibility and evidence intact across the boundaries—not giving every agent broad access to make the hand-offs easier.

In a mid-sized business, several responsibilities may be shared across a smaller team. In a small business, the owner may design the service, approve the proposal and deliver much of it. The paperwork can be lighter, but the difference between what was requested, what can be offered and what was approved still matters.

One person checking their own work is not independent assurance. Where a decision genuinely requires independence, the business needs an appropriate separate reviewer or a narrower operating scope—not a new label for the same person.

For any size of organisation, I would start with four questions: what may this worker do without asking, what must it bring to a person, who owns an unresolved case, and what happens when that person or the digital service is unavailable?

The answer may be less autonomy for some work. That is not a failure to embrace AI. It can be the honest operating choice when the organisation cannot yet support a more consequential delegation.

Build on the disciplines that already exist

This work should sit alongside established management and assurance practices, not pretend to replace them. ISO/IEC 42001 addresses establishing and improving an AI management system. NIST's AI Risk Management Framework provides a voluntary basis for considering AI risks and trustworthiness across the lifecycle.

Those are foundations for the wider discussion. Referencing them does not make my operating model conformant or certified, and neither establishes that the fictional response controls work in a real system. The ISO reference here is to its public summary, not a clause-level assessment of the licensed standard.

What I am trying to make explicit is the management decision between a capable system and an authorised, evidenced outcome. Confidence, trust and assurance help me ask different questions about that decision. They are not substitutes for engineering, human accountability or testing.

The next step is to connect these decisions across the full operating model: how work is commissioned, staffed, governed, reviewed and improved. That is the purpose of the forthcoming Beyond Agents reference paper.

There is more to say about how we induct, develop and review a digital worker; that belongs in the Agent HR article. The detailed economics, agent-to-agent delegation and controlled improvement each deserve their own treatment too. Developing and maintaining a standard offering is also a larger service-design discussion. Here, the offering is our evidence source, not a second operating model squeezed into the article.

For now, think about the next solution or customer response you would trust an agent to help develop. Where would you let it analyse and draft? Which decisions would you expect it to bring to a colleague? What would convince you that the final answer describes the service the team is actually prepared and authorised to offer?

If the answer is only “we trust the agent”, there is still an operating decision to make.

Related reading

AI Is Leaving the Screen. What Happens When It Comes to Work? →

Ghost in the Machine: Lessons from Our Workforce Operating Model →