I love a good proof of concept. It turns a conversation into something you can touch, challenge and learn from. The trouble starts when a successful demonstration is treated as evidence that a dependable product already exists.
A POC answers a deliberately narrow question: can this capability work under the conditions we have chosen? A product has to answer a much less forgiving question: can people depend on it across ordinary days, ugly incidents, changing demand and the slow accumulation of technical and organisational debt?
Those are not different levels of polish. They are different categories of evidence.
This is my working position, formed across infrastructure, product architecture, innovation and managed-service environments. It is not a universal maturity standard. Different products carry different risks. The point is to make the claim being tested explicit, then resist quietly expanding that claim after the demo has gone well.
The method below is a decision aid, not proof of a named deployment. It comes from a hands-on habit: build close enough to understand the capability, keep the evidence visible, then follow the idea into ownership, support, adoption and economics before calling it ready.
The demo has a protected childhood
A useful POC removes variables on purpose. It may have a small dataset, a friendly user, a known workflow and an engineer standing close enough to restart the thing when it sulks. Cost optimisation can wait. Support is a direct message. Security exceptions are temporary. Nobody has yet tried to use it at quarter end, from a slow connection, after an upstream API changed without warning.
None of that makes the POC dishonest. Controlling the conditions is how an experiment isolates a question. If the question is whether a model can classify a certain document, adding a complete service-management layer before testing the capability would usually be premature.
The mistake is forgetting which conditions were controlled. “The capability worked in our experiment” becomes “the solution works.” Then “the solution works” becomes “we are ready to scale.” Each sentence sounds like a modest edit. Each one makes a materially larger claim.
The job of a POC is to reduce uncertainty—not to make uncertainty disappear from the slide.
Production readiness requires surrounding evidence
The prototype often receives most of the attention because it is visible. A dependable product or service includes the system around the capability: identity, data, interfaces, monitoring, support, commercial constraints, change processes, user behaviour and accountable ownership.
A model may produce an impressive answer while the service cannot explain which data it used. An agent may complete a workflow while nobody can show which permissions it exercised. A spatial experience may delight in a controlled room while the device estate, accessibility requirements and support model make broad adoption impractical.
In each case, the capability evidence may be sound. What is missing is evidence about the operating environment around the capability.
- Capability: does the core technology perform the task well enough under stated conditions?
- Context: does it still work with real data, users, dependencies, latency and edge cases?
- Control: can identity, permission, privacy, security and policy be enforced and inspected?
- Operation: can teams monitor, support, recover, change and retire it?
- Economics: do cost, demand, licensing and service obligations remain viable at the expected scale?
- Adoption: does it improve the work, and can people understand, challenge and use it safely?
Name the state before promoting the claim
A surprising amount of confusion sits inside ordinary verbs. Proposed, documented, built, tested, enforced and measured can describe the same idea at very different levels of evidence. I want the state written beside the claim before a team makes the next decision.
- Proposed: the idea or control has been described, but not yet implemented.
- Documented: the intended design or process is written down; that does not mean it is operating.
- Built: an implementation exists; testing and dependable operation are still separate claims.
- Tested: the implementation was exercised within a named environment, date and boundary.
- Enforced: a control demonstrably constrained or prevented an action within the tested boundary.
- Measured: a result has a method, period, unit and stated limitation.
These labels are not an automatic maturity ladder. A team may decide that a documented manual stop is proportionate for a reversible experiment, while a consequential service needs a tested and enforced control. The useful outcome is an honest decision, not the appearance that every box is green.
A product needs an owner on a bad day
“Who owns it?” sounds like a governance question until something fails. Then it becomes an engineering question, a customer question and sometimes a financial question at the same time.
Ownership is more than the name of an executive sponsor. I want to know who makes the next decision, who operates and supports the service, who can accept risk, and who can stop it when continuing would be worse. In a small experiment several roles may sit with one person. They still need to be named rather than left with whoever happens to be online.
This is one reason innovation teams can become trapped by their own success. They prove a capability, demand arrives, and the small team quietly becomes an unofficial production operation. The organisation gets the appearance of speed while inheriting a service with unclear accountability and a permanently heroic support model.
The right response is not to bury every experiment in heavyweight process. It is to decide deliberately when the work changes state. A prototype may remain a prototype. A promising experiment may enter a productisation phase. A narrowly bounded internal tool may require fewer controls than a customer-facing service. The transition should be visible, funded and owned.
The economics arrive late and stay forever
POCs are often priced for learning. Products are priced for living. That means normal demand, redundancy, observability, support time, data movement, model or platform consumption, security controls, compliance work, updates and eventual retirement.
Unit economics also change behaviour. A capability that looks sensible at low volume may become expensive when demand, assurance and support grow. A low per-use cost may hide the human effort required to review uncertain results. An attractive starting price may also weaken the organisation’s ability to change direction later.
I do not believe every number must be known before proceeding. Emerging technology rarely grants that certainty. I do believe the unknowns should be named and assigned. “We have not measured recovery time” is a manageable research gap. Treating recovery as somebody else’s future problem is not.
A practical gate before “scale”
Before moving a POC into productisation, I want a team to be able to answer the following questions in plain language. The answers do not all need to be positive. They need to be honest enough to support a decision.
- What exactly did the POC prove, and under which controlled conditions?
- Which important claims—scale, security, recovery, adoption or economics—remain untested?
- Who makes the next decision, who operates and supports the service, and who has authority to stop it?
- What data, identity and dependency boundaries does it cross?
- What does failure look like, how is it detected and how is service restored or safely stopped?
- How will the capability change real work, and can the people affected understand, challenge and adopt it safely?
- What is the smallest useful production boundary rather than the largest imaginable launch?
- Which costs grow with usage, assurance, support and human review?
- What evidence would make us pause, redesign or end the work?
Productisation should preserve the learning
There is a temptation to treat the POC as disposable and begin the “real” project with a clean slide. That can throw away the most valuable thing the experiment produced: evidence about assumptions, failure modes and user behaviour.
I would rather preserve a short claim ledger. What did we expect? What did we observe? What changed our mind? Which conditions have not been tested? Product architecture can then respond to what was learned instead of reconstructing the experiment as a success story.
The blank ledger below is deliberately small. Use one copy for each material claim. Empty fields are visible research work, not a reason to fill the page with confident language.
- Decision in view: proceed / narrow / redesign / pause / stop.
- Claim: the exact statement the evidence needs to support.
- State: proposed / documented / built / tested / enforced / measured.
- Evidence receipt: source, owner, date, environment and result.
- Boundary: controlled conditions, exclusions and important claims still untested.
- Ownership: decision owner, operator, support path, risk acceptance and stop authority.
- Operating consequence: adoption, data, security, recovery, economics and human review.
- Next evidence: what must be learned, who owns it and when the decision will be revisited.
The productisation phase should also be allowed to reject the idea. That is not failed innovation. Discovering that the operating cost, risk or adoption burden outweighs the benefit is a useful outcome—especially if it happens before customers or teams become dependent on the service.
The gate should change the next decision
This method does not produce a magic readiness score. It makes the decision, evidence and gaps visible to the people who must carry the consequence. Too little discipline turns prototypes into accidental production. Too much asks an experiment to carry controls intended for a mature, high-impact service. Risk, reversibility, audience and blast radius all need to shape the path.
AI makes this calibration harder because behaviour can vary even when the surrounding software has not changed. Spatial and edge systems add device, environment and connectivity constraints. Managed services add obligations that may outlive the technology fashion that created the demand. I am developing a more explicit readiness model across those contexts, but I do not think one universal score will be the answer.
My current test is simpler: have we produced enough evidence for a named person to make a responsible decision about the next boundary? If yes, the POC has done its job. The next job is not “make the demo bigger.” It is to build the product, service and operating model that the promise now requires.
The workbench remains essential. It is where the capability earns attention. Production is where the whole system earns trust.
This note deliberately publishes a working method and a blank artefact. It is not a disguised customer case study, evidence of a named deployment or a claim that one model fits every product. The method should be challenged, adapted and judged by the decisions and evidence it helps make visible.
For the broader pattern behind this position, read From racks to agents, or explore the evidence boundaries on the Workbench.