Why AI pilots stall before production, and what to check first
A pilot answers “can this work?”. Production asks “will it keep working when nobody is watching, and what happens when it’s wrong?”. Most stalled pilots were never set up to answer the second question.

The short answer
Most AI pilots stall because they were built to show that something can work, not to be judged and run. Nobody agreed what a wrong answer costs, there is no test set of real cases, review and fallback steps were never designed, and running cost wasn’t measured. Check those first, before changing the model or prompt.
A typical stalled pilot looks like this. A prototype sorts incoming documents, drafts replies or answers questions from internal files. The demo goes well and stakeholders like it. Then months pass and it still hasn’t gone live. The usual suspicion is the model. It is rarely the first thing to check.
This article explains why pilots stall, sets out the checks to run on a pilot you already have, in order, and describes the three honest ways forward. It follows the same order we use when we take AI from prototype to production.
A pilot and a production system answer different questions
The first is about possibility; the second is about operation, unattended, with consequences. In practice the two differ in four ways.
- Inputs. A pilot runs on examples someone chose. Production gets everything: scanned pages, half-finished emails, the supplier who formats everything differently.
- Audience. A demo audience is forgiving and wants it to work. Daily users are busy, and they quietly stop using a tool that wastes their time.
- Supervision. In a pilot, the builder is watching and fixes things on the spot. In production, nobody is watching unless someone designed a way to.
- Consequences. A pilot often suggests and nothing happens. A production system updates records, sends messages and changes what people do next.
None of these gaps is mainly about how capable the model is. They are about decisions that nobody made while the pilot was being built.
Why pilots stall
Each of the causes below is a missing decision rather than a missing feature.
Nobody agreed what a wrong answer costs
A wrong answer might mean a quick correction, a delayed order or a problem with a regulator. That cost decides how much review the system needs, and if nobody has written it down, there is no basis for deciding. The pilot ends up judged by anecdote: a senior person tries three examples, one is wrong, and confidence goes. The system may be good enough for the job, or it may not. Nobody can say, because nobody agreed what “good enough” meant.
There’s no agreed way to tell whether it works
Without a set of real cases with agreed correct answers, every change is a matter of opinion. A prompt tweak fixes the example someone complained about and breaks two others nobody checked. A newer model “feels better”. The team goes round in circles because there is no score to move. It is also the gap to close first.
The pilot automated the decision, not the preparation
Most of the time saved by automation sits in preparation: reading, extracting, matching and drafting. Most of the risk sits in decisions. Pilots often aim straight at the decision because it makes the best demo, and then nobody is willing to let it run unsupervised. We cover the distinction in more detail in Choosing what to automate.
The human step was an afterthought
“A person will check it” is not a design. A long approval queue with little context turns review into a click, not a decision. If the reviewer can’t see the original input next to the output, can’t see what the system was unsure about, and finds correcting harder than accepting, the review step adds delay without adding safety. People notice, and trust drains away.
The production work was invisible
A demo hides most of what production needs: connections to the real systems, permissions, data handling, running cost, monitoring and a plan for when the model is unsure, slow or unavailable. None of it shows on screen, so none of it was estimated. When it surfaces, the pilot looks much further from done than anyone expected, and the project loses momentum. A working screen is not a working system.
Why AI agents fail in production
Agents stall for the same reasons, made sharper because they act rather than suggest. An agent that updates records or sends messages needs a defined set of tools, permissions limited to what the task requires, approval for steps that carry risk and a full log of what it did and why. Without those, nobody responsible for the process will let it near real work, and they are right not to.
What to check first
If you have a stalled pilot, run these checks in order, before you change the model or the prompt. Each one leaves something written down. It is the same order we follow on new work.
- Name the decision, the owner and the cost of a wrong answer. Which task or decision does the AI support, who is accountable for it, and what does a wrong answer cost: a quick correction, a delayed order, or a problem with a regulator? You should end up with: scope, owner and cost of error, in writing.
- Check that AI is the right tool for this step. If the rules can be written down, plain code or a workflow tool is cheaper and can be tested completely. If errors are unacceptable and nobody can check the output in time, AI shouldn’t take that step. If there are no real cases yet, start capturing them first (where AI isn’t the answer). You should end up with: a yes, a no, or “rules for this part”.
- Separate preparation from decisions. Mark each step as mechanical, preparation or decision. Automate preparation; keep decisions with the accountable person until there is evidence they can move. You should end up with: a step map with the boundary marked.
- Build an evaluation set and run the existing pilot against it before changing anything. Use real cases from your work, including the awkward ones, agree the correct answers with the team, and write a checker that scores them. You should end up with: an evaluation set and baseline results, which show what to keep, change or add.
- Design the approval points and fallbacks. Decide which steps a person approves, what the reviewer sees and how easy it is to correct. Decide what happens when the model is unsure, slow or unavailable: send it to a person, fall back to a rule, or stop. You should end up with: approval points and fallback paths.
- Measure cost per task and response time on the evaluation set. This tells you what a typical month should cost before launch, and where a cheaper model or fewer calls hold quality. If failed tasks have to be redone, count cost per successful task, not per attempt. You should end up with: expected cost per task and a budget with alerts.
- Pin the model version and decide how changes go live. Providers update and retire models. A new model or prompt goes live only after it passes the evaluation set, and the previous version stays available until the switch is confirmed. You should end up with: a change log and a rule for model updates.
- Plan monitoring and a runbook before go-live. Track quality through sampled checks and reviewer corrections, along with cost per task, response times and failures, with alerts. You should end up with: dashboards, alerts and a runbook covering quality drops, cost rises, provider failure and model change.
Don’t start by upgrading the model. A newer model changes behaviour in ways you can only see with an evaluation set. Without one, you can’t tell whether it helped.
Together, those artefacts are a fair description of what “production-ready” means: a system handling a defined part of real work, with its evaluation set and results, decision records, runbook and monitoring, that the people running it understand. One more practical point: a production system should run in your own cloud account, with data leaving it only for model calls you have approved, under business terms that exclude training on your data. Settle that early, because it shapes the rest.
How we test, and why plausible isn’t the same as correct
A convincing demo and a working system can look identical until you score them. Our Agent Skills note shows this clearly, though it was a test of coding agents rather than of business workflows, so treat it as an illustration of method, not as evidence about pilots in general.
In a paired test in one repository, a short, explicit skill file took Claude Sonnet 5.5 from 10 of 30 to 30 of 30 release runs that followed every house rule, and Claude Opus 5.5 from 6 of 24 to 24 of 24. A skill that Sonnet wrote for itself, by studying the repository, looked tidy and passed 0 of 30 release runs. Every run was scored by an automated checker, and the question, conditions and analysis were written down before the runs started.
Read by a person, the self-written skill was plausible. Only a checker running on real cases showed it was wrong every time. Stalled pilots need the same discipline, which is why we measure rather than assume.
Three ways forward
Once the checks are done, the evidence usually points to one of three outcomes.
- Narrow it and ship it. Pick the slice of real work that the evaluation shows the system handles well, and put review on the rest. A first version that handles a narrow slice of real work well is worth more than a broad demo, and it starts collecting the corrections that show where to widen next.
- Rebuild the parts that can’t be fixed. Some gaps are structural: the pilot was never connected to real systems, has no place for a review step, or can’t run in your environment. Keep the evaluation set and what you learned, and rebuild those parts properly rather than patching around them.
- Stop, or replace it with rules. Stop if the checks show the rules are clear enough for plain code, if errors can’t be reviewed in time, or if there are no real cases to test against. Sometimes the right answer is no AI at all.
Stopping on evidence is a result, not a failure, and it is cheaper to learn early than late.
A note on data and accountability
Where decisions significantly affect individuals, UK and EU data protection rules set limits and safeguards for decisions made solely by automated means. A documented human decision point makes those obligations easier to meet and easier to explain. The ICO’s guidance on automated decision-making is a good starting point. This is orientation, not legal advice.
Questions
Why do most AI agents fail in production?
They act rather than suggest, so the missing decisions hurt more: unclear permissions, no approval for risky steps and no record of what they did and why. Give an agent a defined set of tools, permissions limited to the task, approval for steps that carry risk and a full action log. Then judge it against real cases like any other system.
Is it true that 80% of AI projects fail?
The widely quoted figures come from different studies that measure different things, and some are predictions rather than measurements. The 80% figure is often traced to a RAND report, but RAND writes “By some estimates, more than 80 percent of AI projects fail” and cites others rather than measuring it. Its own contribution is five root causes drawn from interviews with experienced practitioners, and it excluded projects that simply prompted pretrained language models (RAND, 2024). What matters more is whether your pilot can be judged and run.
How do we know if our pilot is ready for production?
It passes an agreed evaluation set of real cases, it has designed approval points and fallbacks, its cost per task is known, the model version is pinned, and monitoring and a runbook exist. If any of those is missing, that is the next piece of work, not a reason to switch models.
Should we switch to a newer model to fix a stalled pilot?
Not before you can measure. Pin the current version, build the evaluation set and record a baseline. Then run the newer model against the same cases and switch only if it passes. Without that comparison, a new model just changes which answers are wrong.
When should we stop a pilot?
Stop when the rules are clear enough for plain code, when errors would do serious harm and can’t be reviewed in time, or when there are no real cases to test against. In the last case, the first job is to start capturing those cases. Stopping on evidence is a useful result.
If your pilot has stalled
If you have a pilot that worked in a demo and hasn’t gone live, tell us about it. You can read more about how we approach applied AI first, or get in touch directly.
