Skip to content

Notes

Why the pilot never becomes a system

On the distance between a demo that impressed the board and something the operations team depends on.

Almost every organisation we speak to has already run an AI pilot. Very few are running one in production. The gap is not model quality — the models are better than the processes they are pointed at. The gap is three pieces of unglamorous work that nobody was assigned.

Nobody owned the integration

A pilot runs on an export. Production runs on the system of record, which has an authentication scheme from another decade, a field that means two different things depending on who filled it in, and no API for the one action that matters. Estimating that work honestly makes a proposal look slower than a competitor's, so it gets estimated dishonestly, and the project stalls exactly there. When we scope, integration is quoted as its own milestone with its own risk, because pretending otherwise only moves the failure later.

Nobody wrote down what correct means

Ask most teams what accuracy their agent needs and you get a number with no test behind it. Correct is not a percentage; it is a set of cases with agreed answers, including the ones where the right behaviour is to refuse. Build that set from history before you build the system, and two things happen: the scope becomes arguable in a useful way, and the day a provider changes a model underneath you, you can re-run it in an afternoon instead of discovering the regression through a complaint.

Nobody designed being wrong

Every pilot demonstrates the success path. Production is defined by the other one. What does the system do at low confidence — guess, escalate, or stop? Who sees the escalation, and how fast? What is reversible and what is not, and is anything irreversible allowed to happen without a person? A system that fails visibly and safely will be trusted with more work than one that is slightly more accurate and fails silently.

The narrow thing that runs beats the broad thing that demos

The most useful first deployment we know of is one case type, one team, one month, with a person still in the loop and a switch that turns it off. It looks unambitious in a steering committee and it is the only version that produces evidence. Breadth is easy to add once one team relies on the thing; credibility, once spent on a pilot that went nowhere, is expensive to get back.

If you have a pilot that stalled, the interesting question is which of these three it died on. It is usually identifiable in one conversation.

Tell us where yours stopped.

Integration, acceptance criteria, or failure design. One message and we will tell you what we would do differently.

Our WhatsApp line is being connected. Email reaches us today.