From Pilot to Production: Why Most AI Projects Get Stuck, and What the Exceptions Do Differently
In January I started building a cloud security SaaS. Solo, without a development team, with a handful of AI coding agents. Five months later it is running in production.
And I now understand from my own experience why exactly this step, from pilot to real operation, is the hardest one for most organisations. The numbers are sobering, and they match what I hear in client conversations.
The Uncomfortable Data
Through the NANDA initiative, MIT analysed 300 public AI projects, 150 interviews, and an employee survey in the report "The GenAI Divide". The finding that made the rounds in the trade press: around 95 percent of enterprise GenAI pilots show no measurable impact on the profit and loss statement. Only about 5 percent make the leap.
The exact figure is disputed, and rightly so. But the direction is consistent across studies. In "State of AI 2025", McKinsey arrives at the same pattern from a different angle: 78 percent of organisations use AI in at least one function, but only 39 percent attribute any EBIT impact to it at all. For most of them, the effect is below 5 percent. 62 percent are experimenting with agents, only 23 percent are scaling them anywhere in the company.
The RAND Corporation looked at failed projects and arrives at a failure rate of a good 80 percent: one third is cancelled, another third delivers no value, and the rest does not justify the cost.
Put differently: almost everyone has started. Almost no one has arrived.
What It Is Not About
The first reaction in many boardrooms is to doubt the model. Wrong model, wrong vendor, too little fine-tuning. The MIT team is clear on this: the model is rarely the bottleneck. The bottleneck is the gap between the tool and the organisation.
Certain patterns keep recurring. Integration into grown legacy systems is underestimated. A pilot that shines on a clean dataset fails against a reality of three source systems, two naming conventions, and an interface that no one has touched since 2018. The output quality that convinces in the demo collapses at volume, because ten hand-picked examples suddenly become ten thousand real cases. The monitoring is missing to even see whether the system still does what it should in operation. And no one is clearly responsible when the pilot is meant to go wide.
Behind this often lies the same data problem. RAND traces a large share of the failed projects back to unresolved data questions, not to the AI itself. Anyone without a solid foundation is building on sand, no matter how good the model is.
And there is a thinking error behind it. A large share of companies treat AI as an IT project, not as a business change. An IT project delivers a system and is then done. A business change alters how people work, and is only done once the new way of working has settled in. That is precisely the difference between the 95 and the 5 percent.
What the Exceptions Do Differently
McKinsey named the single strongest factor correlating with real EBIT impact: fundamentally rethinking workflows. It is not the best model that wins, but the organisation that rebuilds its process around the AI, instead of gluing the AI onto the old process.
An example from my own work. The old order was: requirement, write code, test, review. With agents I reversed the order. First formulate the specification cleanly, then let the agent build, then put most of my time into the review. This is not the same process with a faster tool. It is a different process.
On top of that, three things stand out in the successful projects. There is clear business ownership, not just a project team in IT. There is a solid data foundation before the first agent starts. And a human stays in the loop wherever decisions have an effect.
One finding from the MIT study deserves particular attention, especially for the Mittelstand: bought-in solutions and partnerships succeeded in around two thirds of cases. Pure in-house builds only about a third as often. The reflex of "we will build it ourselves" is understandable, but statistically it is the more expensive path. Anyone without their own AI experience in-house buys, together with a partner, not just technology but the detour around the typical mistakes.
What I Learned Myself
Two moments from my own project taught me more than any report.
The first was a positive surprise, right at the start. As a product owner I spent years translating business requirements into user stories and sorting out the dependencies before a single line of code existed. The agent did exactly that in minutes, and the dependency split was clean. I can still do this work myself, and that is precisely why I could tell that it was right. Anyone who has never done the process by hand cannot judge the AI's result. That is not nostalgia, that is quality control.
The second moment was a failure. I accidentally let a plan run unsupervised. During deployment I discovered that the agent had simply removed requirements it had implemented earlier. Working features, just gone. A rollback was out of the question, because too much code had been written by then, and when you sit at the token limit every week, tokens are budget. So I rebuilt the original feature from scratch. That cost time I had not planned for.
The lesson is the same one that sits behind the entire adoption gap. The AI was brilliant at the work I used to do by hand. It was dangerous precisely where no human was looking any more. You can automate the building. You have to own the review.
Four Questions Before You Scale
Anyone looking to lift a pilot into production can spare themselves a lot of frustration if four questions are settled beforehand.
Who is responsible on the business side, not just in the project but in operation? If the answer is "IT", the business connection is missing, and that is exactly where most fail.
Is the data foundation in place? Before an agent works at scale, it must be clear which data it accesses, at what quality, and with what authorisation.
Where does the human stay in the loop? Not everywhere, that would be the old process. But at every point where a decision has an outward effect, a review belongs there.
And finally: are you building yourselves, or with a partner who has already walked the path? The statistics here are clearer than pride sometimes allows.
What This Means for You
Anyone wanting to move AI from pilot to production is not dealing with a tooling problem, but with a governance problem. Clear ownership, a data foundation, guardrails, and genuine review capacity: these are the same disciplines that make up a clean cloud transformation. If you can do cloud governance, you can do AI governance. Ignore both, and you end up in the 95 percent.
This gap between "works in the pilot" and "holds up in operation" is exactly what we look at at Tallence Cloud. Not with yet another tool, but with the question of what the process around the AI has to look like so that it stays productive and auditable.
One question to close with, one that occupies me too: where in your organisation have agents already taken over tasks that no one systematically reviews any more?