Skip to content
ENع
Schedule call
All insights
Field notesAug 2026 · 6 min

Why 95% of AI pilots never reach production

The failure is almost never the model. It’s four things nobody scoped, and you can usually spot all four in the first week.

Momentem engineeringDubai
A dark office with the monitors still lit
In this article
1 · The number itself2 · The four things nobody scoped3 · What we do differently4 · How to tell in week two

Every enterprise we speak to has already run an AI pilot. Most of them cannot tell you what it changed. That is not a failure of ambition or of budget. It is a failure of scope, and the same four gaps show up almost every time.

1 · The number itself

MIT’s NANDA research put it bluntly: around 95% of enterprise AI pilots never make it into production. The same body of work found that externally-built systems succeed at roughly twice the rate of internal-only builds, not because internal teams are weaker, but because they are usually building alongside a day job, without the licence to change the process the AI touches.

Locally the tension is sharper. Cisco’s readiness work puts UAE intent to deploy AI agents at 92%, with 17% reporting the mature processes needed to run them. Ambition is not the constraint.

A pilot proves the model can do it. It does not prove your operation can absorb it.

2 · The four things nobody scoped

When we audit a stalled pilot, the cause is almost always one of these, and often all four at once.

  • ·The exceptions. The demo ran on clean records. The business runs on the 12% that nobody documented: the client with two account codes, the order that skips approval, the field that means something different in two systems.
  • ·The write-back. The output landed in a spreadsheet instead of the system of record, so someone still had to re-type it. Nothing downstream got faster, so nothing showed up in the numbers.
  • ·The approval chain. Nobody decided who signs off what. So either everything needed a human, removing the gain, or nothing did, and finance stopped trusting the output.
  • ·The 7am person. The operator was never in the room. The interface assumed they would change how they work, and they reasonably declined.

3 · What we do differently

None of these are model problems, which is why more capable models have not fixed them. They are scoping and integration problems, so we front-load exactly that work.

  • ·Two weeks with the people doing the work before anyone writes production code: where the money leaks, which decisions are genuinely hard, what the systems can and cannot do.
  • ·A scored eval set built from your own history, not a vendor benchmark, and handed over so your team can rerun it on every change.
  • ·Write-back into the system of record from the first working slice, with a documented rollback path for every automated action.
  • ·Approval gates designed with the operator, on anything that touches money or a customer.

4 · How to tell in week two

The useful signal is not accuracy. It is whether a real record has travelled the whole path inside the first two weeks: in from the source, decided, written back, and visible to the person who owns the outcome. If it has not, the pilot is measuring the model rather than the operation, and the number at the end will not move.

That is also the honest moment to stop. An Audit that concludes “fix the data pipeline first, then revisit” has done its job. Deploying something you will switch off in a quarter has not.

Keep reading

Server racks lit from within
Jul 2026 · 8 min

Running AI on-premise in the Gulf

A hand holding a single marked sheet of paper
Jul 2026 · 5 min

What a two-week AI audit should produce

A dreamcatcher against a bright sky
Jun 2026 · 7 min

Evals your operators can run themselves

Think your pilot is stuck on one of these four?

Forty-five minutes with the people who’d build it. We’ll tell you which one it is.

Schedule call
Tell us the number you want to move.Schedule call