◆ Build note

The $15 permit fee

How a production AI pipeline learns to refuse an answer that is literally on the page, and why every check we run exists because something got through without it.

A customer opened their compliance page and saw that the city's short-term rental permit cost $15. The real fee was $313. The $15 was a late-processing charge that sat two sentences below the fee on the city's own page.

What made this one worth writing up is that nothing had gone wrong in the way you would expect. The system that published the number was not guessing. It had fetched the city's own page, extracted a value together with the exact sentence it came from, and run that pair through a verification gate before publishing. Every check passed.

Right URL, right topic, number literally in the quote, and the answer was still wrong. The kind of wrong that gets a host fined, and the kind of wrong a demo never shows you, because demo data does not contain late charges.

Why the obvious fix is the wrong one

The tempting move is to ask the model to be more careful. Add a line to the prompt: "make sure you pick the actual fee." That fails for a reason worth being precise about. A prompt instruction is a request, and you cannot audit a request. You cannot run it over everything you have already published, and you cannot tell, six weeks later, whether it is still being honoured after a model update. It makes the next demo look better and changes nothing about what the system can prove.

What we did instead was add a check. Not a wish, a predicate: a small deterministic function that takes the quote and the value and returns a verdict with a reason. The rule that caught this case is called fee_quote_ambiguous, and it says: if the quote offers more than one dollar amount, a single value from it does not prove which one is the fee, so refuse, unless the value is the largest amount and the others add up to it (a total with its components), or it leads the quote with a renewal fee after it.

quote "The annual permit fee shall be $313.00. Applications received after March 1 are subject to a late processing charge of $15.00." value $15.00 check two amounts in quote, value is not the largest, not a component sum verdict rejected · fee_quote_ambiguous value $313.00 check two amounts in quote, value is the largest, other is not a component verdict rejected · fee_quote_ambiguous (a human confirms this one, once)

Notice that the rule also refuses $313. That is deliberate. The system does not know that $313 is right; it knows that a sentence with two amounts in it does not settle the question, and it says so. A person reads that page once, confirms the fee, and the confirmation is stored with the same receipt as everything else. The gate's job is not to be clever. Its job is to never publish something it cannot defend.

The pattern underneath

The $15 was one bug. The pattern it taught us is what we now build into every AI system that has to be right about the world, and it has five parts.

1. Every published fact carries a receipt

Not "the model said so." The URL it was read from, the exact quote, the value, the date, and the verdict of every check it passed. A fact without a receipt is an opinion with good formatting. The receipt is what lets a customer click through and see for themselves, and it is what lets us do everything below.

2. A check is a function, not a sentence in a prompt

Each check is deterministic and takes the receipt as input. That means it can be tested with fixtures, it gives the same verdict tomorrow as today, and its verdict comes with a named reason. When someone asks why a number was rejected, the answer is a word, not a shrug.

3. A tighter gate must reach what is already published

This is the one most teams miss. When we added the ambiguity check, it protected every future extraction and did nothing for the facts already on customer pages. So each check is also a sweep: it re-runs over every receipt in the book and quietly un-publishes what no longer passes. The first time we ran one of these it found facts from another government's website sitting on pages for months, and every one of them had passed the checks that existed when it was published.

4. A field that never passes is a bug in the check

If a kind of fact is being extracted thousands of times and promoted zero times, the world is not empty. The check is wrong. We found two fields that had consumed thousands of model calls and produced nothing, and in both cases the check was demanding something the source could never say in those words. Watching the rejection reasons, not just the pass rate, is what surfaces this.

5. Absence is not a fact, and it must never render as a good one

When the gate has proved nothing, the honest answer is "we do not know yet," and the page has to say that. The failure mode to fear is the reassuring default: a requirement that renders as "not required" because nobody proved it was. That is the same overclaim as the $15, pointed the other way.

Getting a model to answer is the easy part. The work is everything after it: proving the answer is real, catching it when it is not, and shipping something a business can safely act on.

What this means for your build

If you are putting AI in front of a decision that costs money when it is wrong, the questions to ask whoever is building it are not about which model. They are: what is the receipt for each answer, what are the checks and can I read them, what happens to answers already out there when a check gets stricter, and what does the system say when it does not know. If those four have crisp answers, the model choice is a detail. If they do not, the demo is the best the system will ever look.

We build the checks into the first version, not after someone gets burned. It is slower for about a week and cheaper for every week after that.

Have a process where a wrong answer costs real money?

Send us a paragraph about it. You get a straight read on what it would take, and a fixed price if it is worth building.

More field notes
What an AI automation actually costsThe model bill, the build, and the part everyone forgets. Why the pilot never shippedFive things that close the gap between a demo and a system.