Why the pilot never shipped
The gap between a demo that works and a system a business can act on. Five things that close it, in the order they usually go wrong.
Every operator we talk to has the same story. Someone on the team, or an agency, or a vendor, built an AI pilot. The demo was genuinely impressive. Everyone agreed it should go live. Then it sat in a Slack channel for four months, and now nobody is quite sure whether it still runs.
This is not a failure of the model, and it is rarely a failure of effort. It is that a demo and a production system are different objects that happen to look alike for the first week. Here are the five places the gap opens, in the order we see them, and what closes each one.
1. The demo ran on the demo data
The pilot read five clean PDFs someone picked. Production is four hundred a month, a third of them scanned, some in the wrong template, a few with a page missing, and one in a format nobody warned you about. The model did not get worse. The inputs got real.
What closes it: build against a sample of the real inputs from day one, including the ugly ones, and treat "getting at the inputs" as a first-class part of the build rather than a detail for later. In our experience it is about a third of the work, and it is the third the demo hides.
2. Nobody wrote down what "right" means
"It works" is a feeling until it is a number. Without a set of real cases with known correct answers, you cannot say how often the system is right, you cannot tell whether a prompt change helped, and you have no way to know when a model update makes things quietly worse. Most stalled pilots stall here: the team can no longer say whether it is good enough, so the decision never gets made.
What closes it: an eval set. Fifty to two hundred real examples, graded, run every time anything changes. It does not need to be fancy. It needs to exist, and it needs to produce a pass rate someone can put in a sentence. Our own verification gate is roughly eighty to ninety percent precise on a hard problem, and we say that number out loud, because a system whose accuracy is unknown cannot be trusted at any accuracy.
3. There is no failure mode
The demo always had an answer, which is exactly the problem. A production system has to be able to say "I do not know," and it has to do something useful when it says it. The most expensive bug in applied AI is not a missing answer, it is a confident wrong one that reads as reassuring. A requirement that renders as "not required" because nobody proved it was required is the same mistake as an invented fee, pointed the other way.
What closes it: a designed "unsure" path. Every output carries a confidence or a verdict, the unsure ones go to a person, and the person's answer is stored so the same case is not asked twice. Absence of an answer is never rendered as a good answer.
4. It was not wired into where the work happens
The pilot lived in a chat window or a notebook. The people whose week it was supposed to save live in the CRM, the ticketing system, the shared inbox, the spreadsheet. If using the tool means going somewhere new, copying something in, and copying the result back out, it will be used twice and then abandoned, no matter how good the answer is.
What closes it: the output lands in the system the person already looks at, ideally without them doing anything. An updated record, a drafted reply waiting in the inbox, a row in the sheet, a ticket with the exception flagged. A chat window is a demo surface, not a workflow.
5. Nobody owned it after launch
Even a pilot that clears the four above will die if it has no owner. Inputs drift. Upstream systems change their exports. Models get replaced. And a scheduled job that runs every night and reports success is not the same as a job that did anything; we have watched a process report green for a week while doing nothing at all, because success was defined as "the job ran" rather than "the work happened."
What closes it: a named person, a runbook, and monitoring that checks the outcome rather than the heartbeat. The eval pass rate, the size of the exception queue, and the bill, looked at weekly. For a single automation this is hours a month. It can be your engineer or it can be us, but it cannot be nobody.
A demo shows you the best the system will ever look. A build is everything that keeps it looking that way on a Tuesday in month six.
The honest version of the pitch
None of this is exotic. It is the ordinary discipline of shipping software, applied to a component that is probabilistic. The reason so many pilots stall is that the demo is so persuasive that the rest gets treated as polish, and then the polish turns out to be most of the work.
When we scope a build, all five of these are in the fixed price. Evals, the failure mode, landing it in your systems, and the handoff with a runbook are not line items you can remove to get a lower number, because a build without them is a pilot, and you already have one of those.
Have a pilot that stalled, or a process that never got one?
Send us a paragraph about it. You get a straight read on what it would take, and a fixed price if it is worth building.