Every enterprise we speak to has already run an AI pilot. Most of them cannot tell you what it changed. That is not a failure of ambition or of budget. It is a failure of scope, and the same four gaps show up almost every time.
1 · The number itself
MIT’s NANDA research put it bluntly: around 95% of enterprise AI pilots never make it into production. The same body of work found that externally-built systems succeed at roughly twice the rate of internal-only builds, not because internal teams are weaker, but because they are usually building alongside a day job, without the licence to change the process the AI touches.
Locally the tension is sharper. Cisco’s readiness work puts UAE intent to deploy AI agents at 92%, with 17% reporting the mature processes needed to run them. Ambition is not the constraint.
A pilot proves the model can do it. It does not prove your operation can absorb it.
2 · The four things nobody scoped
When we audit a stalled pilot, the cause is almost always one of these, and often all four at once.
- ·The exceptions. The demo ran on clean records. The business runs on the 12% that nobody documented: the client with two account codes, the order that skips approval, the field that means something different in two systems.
- ·The write-back. The output landed in a spreadsheet instead of the system of record, so someone still had to re-type it. Nothing downstream got faster, so nothing showed up in the numbers.
- ·The approval chain. Nobody decided who signs off what. So either everything needed a human, removing the gain, or nothing did, and finance stopped trusting the output.
- ·The 7am person. The operator was never in the room. The interface assumed they would change how they work, and they reasonably declined.
3 · What we do differently
None of these are model problems, which is why more capable models have not fixed them. They are scoping and integration problems, so we front-load exactly that work.
- ·Two weeks with the people doing the work before anyone writes production code: where the money leaks, which decisions are genuinely hard, what the systems can and cannot do.
- ·A scored eval set built from your own history, not a vendor benchmark, and handed over so your team can rerun it on every change.
- ·Write-back into the system of record from the first working slice, with a documented rollback path for every automated action.
- ·Approval gates designed with the operator, on anything that touches money or a customer.
4 · How to tell in week two
The useful signal is not accuracy. It is whether a real record has travelled the whole path inside the first two weeks: in from the source, decided, written back, and visible to the person who owns the outcome. If it has not, the pilot is measuring the model rather than the operation, and the number at the end will not move.
That is also the honest moment to stop. An Audit that concludes “fix the data pipeline first, then revisit” has done its job. Deploying something you will switch off in a quarter has not.






