Sapio — data-driven AI
ENRO
C1

From AI pilot to production: why most never pay back

By Vlad TudorLast updated: October 2026Citește în română

Most AI initiatives never pay back because the pilot proves the model works and nobody owns the rest: integration into the real systems and the change in how people work. Only 25% delivered the expected ROI (IBM). What changes that: a named owner, integration built during the pilot, and adoption measured as the success metric.

A pilot that works and then never becomes part of daily work is the most common outcome of AI projects in companies. It is also the most expensive, because the money spent is real and the return stays on paper.

This article is about the step after the pilot. For how to design the pilot itself (scope, metric, stop criterion), see how to run an AI pilot. For the wider list of reasons AI disappoints, Vlad Tudor has written why is AI not delivering results.

Why do most AI initiatives never pay back?

Because the pilot proves the model works, and nobody owns the rest. Only 25% of AI initiatives have delivered the expected ROI over the last few years, and only 16% have scaled enterprise-wide, according to IBM's study of 2,000 CEOs in 33 countries (IBM Institute for Business Value, May 2025).

Other large surveys show the same pattern. Only 25% of Deloitte's respondents have moved 40% or more of their AI pilots into production (Deloitte, State of AI in the Enterprise 2026, January 2026). In BCG's survey, 72% of executives named integrating AI with existing systems, tools and APIs as a challenge, and 77% named people adapting to the change and using AI every day (BCG, The Widening AI Value Gap, September 2025). BCG puts 70% of AI value in people and processes, 20% in data and technology, and only 10% in algorithms.

The practical conclusion: the model is rarely the problem. The work that decides whether the investment comes back is the integration into your systems and the change in how people work. In most pilots, nobody is assigned to either.

What is the difference between a pilot that works and a system in production?

A pilot can be a technical success and still be nowhere near ready for production. The differences show up on seven lines:

PilotProduction
DataAn export, a clean sampleThe real daily feed, with everything ugly in it
UsersA few keen volunteersEveryone in the process, sceptics included
ExceptionsSkipped, or handled by hand by the project teamHandled by the system or routed clearly to a person
IntegrationCopy and paste, a separate screenWrites straight into the ERP, CRM or invoicing system
OwnerThe project teamA named business owner and a named technical owner
What you measureModel accuracyThe process outcome, and how much the system is used
CostIgnoredCost per case, tracked monthly

Every line in the third column is work. If it is not planned from the start, the pilot stays a pilot.

Who should own an AI system after the pilot?

Two people, named before the pilot starts. A business owner, usually the head of the process, who answers for adoption and the numbers. And a technical owner, who keeps the system running, updates the evaluation set and decides what changes. If either is missing, the system degrades without anyone noticing.

Klarna is the public example of a system that hit its volume target and missed its quality one. In 2024 the company announced that its AI assistant was doing the work of hundreds of support agents. By May 2025 its CEO was saying "we went too far" and that the result had been lower quality, and the company began hiring people again (Forbes). The launch worked. What happened in the months after it is exactly the part somebody needed to own.

Why does integration decide whether a pilot reaches production?

Most pilots run next to the process: a separate screen, an exported file, a person copying the result into the real system. That is fine for a test. It is not fine for daily use, because every extra step is a reason to go back to the old way of working.

  • Build the integration during the pilot, not after it. At least one real write-back into the core system, under real access rights.
  • Handle exceptions from the start: what the system does when it is unsure, and who it sends the case to.
  • Put logging and cost monitoring in place before the first real user, not after the first surprise invoice.
  • Test on the real data feed, with its formats and errors, not on the export cleaned up for the demo.

In practice, integration often means an older ERP with no API, an accounting package and a national e-invoicing system. That is where the months go, and a pilot that avoids them tells you nothing about production. What we integrate is described on our AI process automation page.

How do you get people to actually use the AI system?

Adoption does not arrive on its own after launch. It is built, with the same discipline as the rest of the system:

  • Involve two or three people who do the work today from week one. They know the exceptions and will be the first users.
  • Measure every week how much the system is used, not only how accurate it is.
  • When someone routes around it, find out why. The reason is usually concrete: a missing field, a case it does not handle, one extra step.
  • Set a date when the old way of working stops for the cases the system handles well.
  • Keep a person in the loop for high-stakes decisions. Trust comes faster when people see they can correct the system.

Training matters too, but on the real process rather than on AI in general. More in how to train your team to use AI.

How do you know whether it paid back?

Only by measuring it on your own process. Before the pilot, write down how the process performs today: time per case, error count, volume, cost. Then measure the same things afterwards, on the same kind of cases. No ROI figure on a vendor's slide, ours included, replaces the one measured on your process. The method is in how to calculate the ROI of an AI project.

What has to be true before you call it "in production"?

  1. It runs on the real data feed, every day, without the project team stepping in.
  2. It writes its result straight into the systems people work in.
  3. It has a business owner and a technical owner, by name.
  4. It has an evaluation set that runs on every change.
  5. It has logging, alerts and cost per case tracked.
  6. The people in the process use it, and the adoption number is measured.
  7. Your team can change it without the provider.

If any point is missing, you have an extended pilot, not a system in production.

This is exactly the part we work on at Sapio, with an embedded AI engineer in your team who takes one use case from pilot to daily use. If you have a pilot that works and nobody uses, start with the Tech Call: 30 minutes, one process, an honest answer. Request it here.

Sources

Only 25% of AI initiatives have delivered the expected ROI and only 16% have scaled enterprise-wide (IBM Institute for Business Value, 2,000 CEOs, 33 countries, May 2025).

Frequently asked questions

Why do AI pilots fail to reach production?

Because the pilot tests the model, and what decides production is something else: integration with the core systems, handling exceptions and getting people to use it. If nobody is assigned to those before the pilot, it ends with a good presentation and no system in use.

What percentage of AI projects deliver ROI?

Few. IBM's study of 2,000 CEOs in 33 countries, published in May 2025, found only 25% of AI initiatives had delivered the expected ROI and only 16% had scaled enterprise-wide. The only figure that matters for you is the one measured on your own process.

Who should own an AI project in a company?

Two named people: a business owner, usually the head of the process, who answers for adoption and the numbers, and a technical owner who keeps the system running. Both are named before the pilot, not after launch.

How long does it take to go from AI pilot to production?

It depends more on integration and data access than on the model. The practical rule: plan the production work into the pilot, not after it. A pilot that has run for months with no production date in writing is the signal that nobody owns the next step.

Should we stop a pilot that works technically but nobody uses?

First find out why nobody uses it. If the reason is concrete, such as an extra step, a missing integration or an unhandled case, fix it. If the process does not actually need the system, stop it. A pilot stopped in time is a good result.

Want to discuss a project?

Book a free discovery call with the Sapio team.

From AI pilot to production: why most never pay back | Sapio AI