Our AI pilot stalled and never went into real use
Short answer
This is not a rare failure: 86–89% of AI agent pilots never reach production, and Gartner expects over 40% of agentic projects to be cancelled through 2027. The top causes are not the model — they are the absence of a way to measure right and wrong, the absence of authority limits, and data that is not ready. All three can be closed in 4–8 weeks.
Typical duration: 4–8 weeks from stalled pilot to production
Review USD 2,500 · execution quoted after the findings
What this looks like
- The demo impresses, but nobody dares use it for real work
- Nobody can answer "what percentage of its answers are correct?"
- Legal or compliance blocks it because there is no decision trail
- The pilot has run for months with no launch date
- Cost per run is unknown, so nobody dares increase volume
Why this happens
The pilot was built to convince, not to run. The difference is large: convincing only needs to work on ten chosen examples, while running has to be defensible across a thousand nobody chose.
So when it is time to switch on, three things that were never built are suddenly missing: a way to measure right and wrong, clear limits of authority, and a decision record that can be inspected. Without those, no leader will sign off.
How we solve it
- 01
Pilot review
We trace what exists, where it stopped, and whether the problem is worth solving with AI at all. One week.
- 02
Measurement
Build a test set from your real work, then measure how often the answers are correct — a number, not an impression.
- 03
Authority limits
Define what the system may decide alone and what must escalate, then enforce it in code.
- 04
Audit trail
Every decision recorded: input, reasoning, output, and who approved it.
- 05
Cost & fallback
Record the cost of every run and add a fallback provider so the service does not die with one vendor.
- 06
Staged launch
Switched on for a small slice of work first, raised once the numbers hold.
Numbers from our own work
86–89% of AI agent pilots never reach production (Forrester, Gartner, McKinsey 2026)
Top blockers: no way to evaluate (64%), governance friction (57%), model reliability (51%)
AI cost can be tiny once measured: rewriting 66 service packages on our own site cost USD 0.06
Mistakes we keep seeing
- Swapping models repeatedly when the problem is the data and the authority limits
- Judging success by user impressions instead of a test set
- Granting full authority on day one, then switching it off permanently after one mistake
- Not recording cost per run, so nobody dares raise the volume
Questions we are asked most
Do we have to throw the pilot away and start over?
Usually not. What is missing is rarely the part already built — it is the measurement, the authority limits, and the audit trail around it.
How do you measure right and wrong?
With a test set drawn from your own work, re-scored on every change. The result is a number you can compare across versions.
Does our data have to leave the company?
Not necessarily. If the data is sensitive, the system can run on your own infrastructure.
What does it cost to run once it is live?
We record the cost of every run from day one, with an alert when it approaches the limit you set.
What if our problem turns out to be a bad fit for AI?
We will say so in the first week and name a cheaper way to get the same result. That serves both sides better than a project that ends in a drawer.
Updated 29 July 2026 · Neuraltan
If this is happening to you
Tell us the situation. We will say plainly whether this is worth doing now, and how long it takes.