What an AI automation actually costs
The model bill, the build, and the part everyone forgets. Real numbers for a first automation, and the one question that decides whether it is worth doing at all.
Most conversations about the cost of AI start with the wrong number. People ask what the model costs, and the answer is that for almost any process inside a normal business, the model is the cheapest part by a wide margin. The money is in the build, and the risk is in the part nobody budgets for.
Here is the whole picture, with the numbers we actually see. Where a figure is illustrative rather than measured, it says so.
1. The model bill
Model pricing is per token, which is roughly per word. The current list prices from Anthropic, per million tokens, are the useful reference points:
What that means in practice. A twenty-page PDF is about fifteen thousand tokens. Reading it with Sonnet 5 and asking for a structured summary costs around three cents. With Opus 5, around eight. Processing a thousand of them a month with the most capable model is a bill of under a hundred dollars.
The measured version, from a system we run: an extraction pipeline that reads government documents and has to be right about them costs around $40 to $60 a day at full speed across more than a thousand jurisdictions, and most of that is speculative research that fills in places nobody has asked about yet. The customer-facing part is a fraction of it.
Two things move this number, and both are engineering choices rather than model choices. Prompt caching cuts the input cost of repeated context by about ninety percent, but only if the prompt is built so the stable part comes first, and we have measured a lane where caching was switched on and cost more than no caching, because it wrote to the cache every run and read from it almost never. And the batch API halves the price for anything that does not need an answer in the next minute, which is most back-office work.
2. The build
This is where the money is, and it is the part a demo hides. Getting a model to produce a good-looking answer on a clean example takes an afternoon. Getting a system that takes real inputs from the three places they actually live, handles the ones that are scanned or half-filled or in the wrong format, does the work, puts the result where the person already looks for it, and tells someone when it is not sure: that is weeks, not hours.
Our fixed prices for this are public. Most first builds land between $15,000 and $40,000 and take three to six weeks. A small, well-understood automation, a single document type into a single system with a clear right answer, can come in under that. Something that is really several builds gets split into phases, each priced on its own, so you can stop after any of them.
What is inside that number, in roughly the order the time goes:
- Getting at the inputs. Access to the systems, the exports, the mailbox, the shared drive. Usually a third of the work, and the third that is invisible in the demo.
- The core. The prompts, the structured outputs, the code around the model that turns "a model answered" into "a record was updated." Often the smallest part.
- Evals. A set of real cases with known right answers, run every time anything changes, so "it works" is a measured pass rate rather than a feeling. Without this, you do not know when a model update quietly makes things worse.
- The failure mode. What the system does when it is not sure. It has to be able to say "I do not know" and route to a person, because absence of an answer is not an answer, and the most expensive bug is a confident wrong one.
- Landing it where the work happens. The output has to arrive in the tool people already use. A chat window is not a workflow.
- Handoff. Documentation, the runbook for when it breaks, and the code in your repository under your name.
3. The part everyone forgets
Nothing that touches the outside world stays correct on its own. Suppliers change their invoice layouts. A system you read from changes its export. The model you built on is replaced by a better one that phrases things differently. Someone in the business starts using a field for a new purpose.
So a running automation needs a small amount of ongoing attention: someone watching the eval pass rate, the exception queue, and the bill, and making the occasional fix. For a single automation this is hours a month, not days. It can be your own engineer with the runbook we hand over, or it can be us. What it cannot be is nobody. An automation that nobody owns is one that will drift until someone notices, and by then it has been quietly wrong for a while.
The other forgotten cost is the human review lane in the first weeks. A new automation should run with a person checking a sample of its output until the eval numbers earn it more trust. Budget a few hours a week of someone's time for the first month. It is cheap, and it is how you find out what the demo did not show you.
4. The one question that decides it
None of the numbers above matter until you know one more: what does the manual work cost you today? Not roughly. Who does it, how many hours a week, at what loaded cost, and what happens when it is late or wrong.
An illustrative case, with round numbers. Two people spend six hours a week each pulling orders out of three systems into a spreadsheet and chasing exceptions. At a loaded cost of $45 an hour that is about $28,000 a year, before counting the mistakes and the meeting that starts late every Monday because the sheet is not ready.
That is a good deal, and it is a typical shape. If the same maths gives you a payback of four years, the honest answer is not to build it, and we will say so on the first call. The AI Opportunity Sprint exists mostly to run this calculation across every candidate process in a business before anyone commits to the first build, because the process people are loudest about is not always the one that pays back first.
What to take from this
- The model is cheap. Do not let anyone anchor the conversation on it.
- The build is the cost, and evals and the failure mode are part of the build, not extras.
- Someone has to own it after launch. Decide who before you start.
- Know what the manual work costs before you price the fix. It is the only number that tells you whether to do this at all.
Want the maths run on your process?
Send us a paragraph about the manual work. You get a straight read on what it would take, and a fixed price if it is worth building.