
AI · Data · 23 August 2026 ·
Why AI initiatives stall in the pilot phase, and how to reach production: ownership, a data foundation that holds, and proper operations.

The demo was convincing. The steering group was enthusiastic. The model did exactly what it was supposed to do. And then: nothing. Six months on, the proof of concept is still running on a data scientist’s laptop, and nobody quite dares to ask when it will actually be put to use.
We see this pattern in our advisory work more often than we would like. In our experience, the move from proof of concept to production is the stage where most AI initiatives get stuck: not the build, not the model, but the step after it. The good news: the causes can be named quite precisely, and some of them you avoid by setting the proof of concept up differently. Below are the reasons we come across most often, and what you can do differently as a decision-maker.
The distinction we work with: a proof of concept tests technical feasibility, a pilot tests whether something holds up in the real working process. Organisations use these terms differently, so agree at the outset what you mean by them: that looseness creates expectations that derail a project before it has started.
A successful proof of concept therefore mainly shows that the technology can work. Not that an investment in production is justified.
Many proofs of concept also start from a technology-first question: “what could we do with generative AI?” As an exploration that is fine. As the basis for a production decision it is not: that calls for a result named in advance: lower costs, time saved, better decisions, higher quality, or making a critical process more scalable.
Without that agreed result, what follows is predictable. The proof of concept “succeeds”, everyone nods, and then no owner frees up any budget. Not because the result disappointed, but because nobody agreed what success means and who is supposed to act on it.
In the boardroom the two terms are used interchangeably, and that costs time later on. The distinction we apply: a proof of concept tests technical feasibility; a pilot runs with real users, real working processes and real operational conditions, in a controlled but production-representative environment: one step away from production.
Treat a successful proof of concept as an almost-finished product and you skip the stage in which many practical problems surface. The risk is that they only become visible after go-live, when correcting course is considerably harder.
A gap that in our experience is often underestimated sits in the data. In many proofs of concept we come across, the data was gathered, cleaned and prepared by hand, once, by someone who knows exactly which odd exceptions need filtering out.
In production that luxury is usually gone. There you want data to flow out of source systems reliably and repeatably. If those flows are brittle, so is the model’s output, however good the model itself may be. So put transformations under version control and make the origin of data traceable, from source system to model output.
Gaps in data governance are among the most stubborn causes of stalled projects we encounter. Before going into production, you want answers to questions such as:
If the answer to these questions is “we will sort that out later”, then later is now.
The pilots that run aground were, in our experience, almost always built as an isolated experiment rather than as part of a system landscape. A prototype can demonstrate feasibility in a sheltered environment, but it lacks the infrastructure to handle real variability, security requirements and operational load.
Production imposes requirements that received less attention in the proof-of-concept phase:
Integration. The model has to talk to existing systems: case management systems, ERP, customer portals. In practice those connections often take more work than expected, certainly when the teams who will run the system are involved late.
Security and privacy. Worth being clear about: the GDPR applies the moment you process personal data, in a proof of concept or a sandbox too. Which obligations apply exactly depends on your role, the purpose, the legal basis and the risk, among other things; a data protection impact assessment is not mandatory in every case. If you are already working with real personal data in the proof-of-concept phase, make sure the basics are in order and take advice if in doubt. Production usually adds stricter operational safeguards on top of that: worked-out authorisations, appropriate logging, and security that stands up to real use at scale. What counts as “appropriate” follows from risk and context, not automatically from the law.
Operations and monitoring. What monitoring and availability you need also follows from risk and business impact: which processes will depend on this system, and what service levels go with them? Who steps in when the system behaves oddly, and how would you notice in the first place? Those questions deserve an answer before go-live.
Cost at scale. Infrastructure and processing costs at high volume can be modelled or tested in advance. Do that before the production decision, not after it, in an uncomfortable conversation with the controller.
Technology is by no means always the only reason AI initiatives stall. There is a reason adoption by the team that has to work with the system is one of the four areas we test before a pilot moves to production, alongside data quality, infrastructure costs and security.
In many proofs of concept we see, the employees whose work will change are largely left out. When the result is then “rolled out”, it meets fair questions: does this fit my job, can I rely on it, what happens when it gets something wrong?
An AI application that touches the working process shifts responsibilities, or at the very least asks that they be set down more explicitly. Who is accountable for a decision based, in part, on model output? As long as that is unclear, there is a real chance that staff will leave the system alone, and that is understandable.
The lesson: involve process owners and end users during the proof of concept, not after it. Run the pilot inside the real working process, with real users, and measure not only whether the model is right but also whether people use it.
A pattern we recognise: the proof of concept ends in a presentation, not in a decision. There is no threshold agreed in advance above which the initiative continues and below which it stops. The result has a name: pilot purgatory: the state in which a working proof of concept never reaches production, but is never formally stopped either.
That does damage. An initiative that drags on keeps demanding attention and can undermine confidence in the ones that follow. In our view an honest “no, we are not going to do this” is worth more than a permanent “maybe”.
A usable decision framework is set down before the proof of concept starts. One form that works well in our experience is financial gating: only scale up once the modelled payback period meets a threshold set in advance, verified with finance and not just with the project team. Translated into four criteria:
The common thread: the gap between proof of concept and production usually opens at the start, not at the transition. A proof of concept designed to impress produces a demo. A proof of concept designed as the first step towards production produces knowledge you can reuse in the next phase.
In practice that means:
Choose a use case with an owner. Not “something with AI”, but a concrete process with a process owner who wants the result and takes responsibility for it.
Test on representative data. Use data as it looks in production, mess included. A model tested only on cleaned data has demonstrated feasibility on tidy data, not that it holds up in the real working process.
Document what is reusable. A good proof of concept or pilot produces artefacts you will need again in the next phase: interface specifications, role and authorisation models, fallback procedures. Record them while you make them, not afterwards from memory.
Work out in advance what production costs. Not to the euro, but the order of magnitude: infrastructure, operations, integration, user training. That reduces the chance of the business case collapsing only after the demo.
Schedule the pilot phase. Explicitly set aside time and budget for a pilot with real users before you decide on production. Skipping that phase increases the risk that problems only come to light after go-live.
One sober note to close on. How long the road from proof of concept to production takes varies a great deal per application, but making data flows robust, building integrations, setting up security and bringing users along is serious work. Do not count on it going quickly by itself, and plan realistically for it from the start.
That is not a reason not to begin. It is a reason to begin small, with a use case that is worth it, and to take the preconditions seriously from day one.
Do you have an AI initiative that has been sitting in the proof-of-concept phase for months? Then do not schedule another demo, but a half-day review session with the process owner, someone responsible for the data, someone from IT operations and someone from finance.
Answer the same four questions together: can the data supply be made production-worthy, what does running at scale cost, does it meet security and privacy requirements, and who is going to use and maintain it? Every question that goes unanswered is your next action.
And be willing to accept the outcome, including “stop”. A decision that is actually made, in whichever direction, brings clarity and frees up attention and capacity for the initiatives that are worth it. If you lack the capacity to run that review yourself, an AI implementation partner who stays on board after the decision helps.
How we make this. AI writes the first version, a second AI checks it, and a person at Twentynext reads and approves every article before it goes live.

Sixteen questions, seven minutes, and an instant spider chart showing your strongest and weakest dimension. No e-mail address needed to see the result.

AI · 21 September 2026
Sixteen questions, seven minutes, four dimensions on five levels, no email needed. What the scan measures, how it compares and how to use it.
Read the article →

AI · 6 September 2026
Built an app with AI and usage is growing? How we take it into managed service: review the code, make it scale and have it pentested independently.
Read the article →

Cases · 6 September 2026
Read the article →
Work with us
The people who build it also run it afterwards. Eindhoven, since 2014.

Martijn van Grieken
Director Data & AI
We use Google Tag Manager to measure visits and Leadinfo to recognise which company is visiting. Neither loads unless you agree. If you choose essential only, the site works as normal and we measure nothing. If you arrived via an advertisement in ChatGPT, we also use the OpenAI measurement pixel to attribute conversions to that advertisement. Cookie statement · Privacy statement