Systems and Empathy

Why do most AI pilots fail?

· 7 min

A small figure standing in a dark room before three tall doorways, each pouring yellow light onto the floor. A decision constrained by structure.

The figure that gets quoted is that roughly 95% of enterprise generative AI pilots produce no measurable P&L impact. It is usually deployed as evidence that the technology is overhyped.

That is the wrong reading. In most of these cases the technology worked. The model did what it was asked. The integration held. The users, when observed, were reasonably impressed.

And nothing changed.

Both of those things are true at once, and the gap between them is not a technology problem. It is a structural one, and it is the single most reliable pattern we see inside large organisations.

What the 95% is actually measuring

It is measuring the distance between capability and consequence.

An organisation that is systems-rich and empathy-poor is very good at acquiring capability. It can buy the tools, deploy the platform, document the process, and produce a slide showing all three. What it cannot do is change what people actually do at eleven o’clock on a Tuesday.

So the capability arrives, sits alongside existing behaviour rather than replacing it, and the P&L never hears about it. The pilot is judged a success by the team that ran it and is quietly not extended.

The tell is always the same: capability everywhere, standards nowhere. Tools in active use across every function, no shared guidance, and no line of sight between what people do internally and what reaches the customer.

Most AI pilots are not pilots

Here is the distinction that resolves most of this, and it is not a semantic one.

A demonstration proves that something can work. It is designed to succeed. Its output is confidence. Its audience is a sponsor who needs to justify the budget already spent.

A pilot changes one real decision, with real data, and is permitted to fail. Its output is a number you did not have before. Its audience is whoever has to commit the money next.

Almost everything labelled an AI pilot in the last three years has been a demonstration. You can tell them apart with three questions.

Demonstration
Pilot
Designed to
Succeed
Answer a question
What it produces
Confidence
A number you did not have before
Who it is for
The sponsor justifying budget already spent
Whoever has to commit the money next
Where it runs
Alongside the real work
Inside it, on a live route or market
Allowed to fail
No
Yes, and cheaply
01Did it change a decision that was going to be made anyway?
02Was there a baseline?
03Was anyone allowed to say it failed?
Three noes means you ran a demonstration. That is not a failure of the technology; it is a failure to attach a consequence.

Did it change a decision that was going to be made anyway?

If the pilot ran alongside the real work rather than inside it, it was a demonstration. Real pilots take something with a consequence attached: a live route, a live product vertical, a live market. Then they change one thing about it.

Not a sandbox. Not a volunteer cohort. Not the enthusiasts.

Was there a baseline?

Most AI pilots cannot state what the thing they improved was worth before they improved it. Without that, “the team found it useful” is the only available finding, and it is not a finding.

A baseline is unglamorous and takes longer than the pilot. It is also the only thing that turns an anecdote into a number.

Was anyone allowed to say it failed?

This is the one that decides it. If the pilot’s sponsor would be professionally damaged by a negative result, you have not built a pilot. You have built a demonstration with a longer runway, and everybody in the room knows it.

The bias case, and why it worked

In 2017 we were asked by EY to make the case for a new global consulting practice around bias in AI systems.

The underlying claim was true and completely abstract: programming has been done predominantly by men since inception, and that imbalance is embedded in systems now running organisations worldwide. Stated that way it sounds like a conspiracy and boards treat it accordingly.

So rather than argue it, we built it. A triangulation sprint with stakeholders, AI specialists and social scientists, then a live demonstration: crowdsourced AI frameworks constructed using the actual gender and ethnic imbalances found in Silicon Valley. The resulting AI had a personality. We unveiled it at Davos and called it #FixTheBug.

What made that work was not the technology, which was modest. It was that an abstract problem became something a boardroom could see, argue with, and act on. It launched a global consulting offer built around diagnostic tools for organisational AI bias.

The lesson generalises. The constraint on AI adoption is almost never the model. It is that nobody senior can see the thing clearly enough to act on it, and a demonstration designed to reassure them does not fix that.

What we get wrong

A practice built on piloting that claims a perfect record is not credible, so here are two of ours.

Fuller. An acceleration and innovation strategy for facilities management. It never took off.

Jadex. The structure held and the model worked. The exit slipped on a management change. The right answer, overtaken by a change in who was in the room.

We keep both visible deliberately, because the alternative claim, that every pilot we have run has scaled, would tell you we are not really piloting.

We pilot so that failure is cheap. Not every pilot scales; that is the point of running one. The alternative is discovering the identical fact after the budget is committed, which costs the same to learn and considerably more to admit.

Against a market where roughly 95% of AI pilots deliver nothing measurable, being able to say that honestly is worth more than a clean record would be.

What to do instead

If you are about to commission an AI pilot, three changes will do more than any change of vendor.

Pick the decision first, not the technology. Start from a commitment you are going to have to make, whether a route, a market, a hire or a spend, and work backwards to what would need to be true. If no decision is waiting, you do not need a pilot yet.

Fund the baseline. Budget for measuring what you have before you measure what you changed. Teams resist this because it is dull and produces no demo. It is the difference between a story and a number.

Name the failure condition in writing, before you start. Agree what result would cause you to stop, and agree who is allowed to declare it without professional cost. If you cannot answer the second half, fix that before you spend anything.

None of this is about AI. It is the ordinary discipline of contained-scale change, applied to a technology that has made it unusually easy to skip.


Systems and Empathy is FastMora’s method: correcting the blend between designable structure and how people actually behave inside it. A pilot is how you find out which half is broken without betting the company.