ARKONE
A long queue of paper trays on an office desk, one tray overflowing

Why 95% of AI Pilots Never Touch the P&L

August 9, 2026 · 5 min read

S
Sobin George Thomas

MIT's NANDA study found 95% of generative AI pilots produce no measurable business return. In the GCC, 63% of organisations are still pre-implementation. The constraint was never the model. It is where the work waits.

Two numbers describe the same gap from different sides. MIT NANDA’s State of AI in Business study reviewed more than 300 enterprise AI initiatives and found roughly 95% produced no measurable effect on the P&L. And Deloitte’s GCC survey of senior finance and tax leaders found more than 63% of organisations in the region still in pre-implementation stages, with only 9% scaling anything.

The comfortable reading is that the technology is not ready. The evidence says otherwise. Read the failures closely and almost none are failures of output. The systems generated the summaries, the drafts, the classifications, and generated them well. They were pointed at the wrong constraint.


The pilots did not fail at output

NANDA’s researchers interviewed 52 organisations and surveyed 153 senior leaders. The pattern behind the 95% was not broken models. It was systems that never touched the number the sponsor cared about. High adoption, low transformation: seven of nine industries showed what the report calls widespread experimentation without structural change.

That distinction matters because it changes who owns the fix. A capability problem belongs to the vendor and resolves in a future model release. A deployment problem belongs to whoever chose what to automate. In 95% of cases, that choice decided the outcome. The model never got a vote.


Automating the step that was already waiting

Consider a pattern recognisable in a dozen companies. A procurement function deploys a model to produce supplier-risk summaries, cutting a two-day analyst task to twenty minutes. The saving is genuine and cleanly measured. Throughput of signed supplier contracts does not move at all, because the binding constraint was a legal review queue running at six weeks.

The organisation automated the step that was already waiting on the step it did not touch. The gain was real, and the queue absorbed all of it.

The practical consequence for any board: stop asking which tasks a model could do. Ask where work currently waits, and for how long. If the model does not touch the longest queue, the business case is an accounting exercise. Measure the wait times before you approve the pilot, not after it disappoints.


When everyone consults the same oracle, errors stop cancelling

A group outperforms its average member only when the members’ errors are independent. Correlate the errors and consensus becomes confidently, collectively wrong, with all the reassurance of agreement and none of its substance.

Enterprise AI correlates errors by design. Most serious deployments sit on a handful of frontier models. Within one company, retrieval runs over one document corpus, through one prompt template, maintained by one platform team. Five analysts consulting that stack are not five analysts. They are one analyst, sampled five times.

The countermeasure is unglamorous and cheap: keep a live control arm. Route 5–10% of consequential decisions through unassisted human judgement, permanently, not as a one-off audit. When the assisted and unassisted populations diverge, you have found either a model failure or a human one. Both are worth knowing, and the divergences are rarely where anyone expected.


An unsigned judgement is not a control

Model output arrives orphaned. A model-drafted paragraph and an argued one look identical in the finished memo. The reviewer cannot tell which sentences someone would defend under questioning and which merely survived a skim. Accountability decays one unattributed paragraph at a time, until nobody in the chain can say who actually believed the number.

Three fields fix most of it, applied to every model-derived judgement that reaches a decision-maker: who is accountable for this being right, what context produced it, and one line on what would change their mind. Then store the judgement rather than regenerating it on demand. A stored judgement has an author and a date. A regenerated one is different every time it is asked, answers to nobody, and bills you again on every viewing.

This is the same discipline that makes an agent auditable in the first place. The case for it is in why AI agents cheat, and the governance that actually holds.


What pre-implementation actually costs in Dubai

The GCC’s 63% is not a technology lag. The region’s leaders expect AI to matter — 93% in the same Deloitte survey predict significant organisational impact. The gap sits between mandate and first working system, and it is exactly where the 95% failure pattern breeds: long evaluations, broad frameworks, pilots scoped to prove a thesis rather than move a queue.

The escape is narrow scope with a measured baseline. One process, one owner, one queue you can time before and after. A system that owns a real piece of work by the end of a day is a data point no strategy deck can argue with, and it is the only kind of evidence that moves an organisation out of the 63%.


Pick the process where work waits longest, and start there. A clinic day takes one of your processes, builds a working agent on it, and leaves you with the measured before-and-after by evening — the anti-pattern to the 95%. If the pilot graveyard is already full, the five-step consulting method starts by finding the queue, not the use case.

Frequently asked questions

Why do 95% of AI pilots fail to show ROI?+

MIT NANDA's 2025 study of over 300 enterprise AI initiatives found the failures were rarely failures of output. The systems generated good summaries, drafts, and classifications. They were pointed at the wrong constraint: automating steps that were already waiting on a slower queue elsewhere, so the gain was absorbed before it reached the P&L.

What should we measure before approving an AI pilot?+

Measure where work waits, and for how long, before approving the pilot, not after it disappoints. If the AI does not touch the longest queue in the process, the business case is an accounting exercise. Wait times, not task times, predict whether throughput will move.

How many GCC organisations are still pre-implementation on AI?+

Deloitte's 2025 survey of senior tax and finance leaders across Saudi Arabia, the UAE, Qatar, and Kuwait found more than 63% of GCC organisations remain in pre-implementation stages. Only 9% have begun scaling solutions.

ഒരു സംഭാഷണം ആരംഭിക്കാൻ തയ്യാറോ?

അളക്കാവുന്ന AI ഫലങ്ങൾ ഡെലിവർ ചെയ്യാൻ ArkOne എങ്ങനെ ഭരണ സംവിധാനവും പ്രോഗ്രാം ആർക്കിടെക്ചറും നിർമ്മിക്കുന്നു എന്ന് കാണൂ.

ഡിസ്‌കവറി കോൾ ബുക്ക് ചെയ്യുക