Today I turned on a check that refuses to save a generated document if it contains an amount, a large number or a clause reference that doesn't appear, character for character, in the source. There's no model in it, just a few regular expressions and a substring search. Within hours it had stopped three runs. One was a real fabrication, the thing it was built for. The other two were my mistakes, and one of those can't be fixed by a better regex at all.

The part of my agent that's allowed to stop a run has no model in it; the model-based judge I also have is only allowed to watch. That only works if you're honest about which claims a dumb check can judge. Figures that should be copied from the source, yes. Figures the workflow is supposed to compute, never. What I'd pass on is less the regex than that split: decide per workflow which kinds of claim are "copied", let a deterministic check block those, and write its rejection for the model, so the model can fix the document itself.

Why the blocking part has no model

My document workflows run as an agent. One model call pulls fields out of the source, a second writes the final document from them, and then the agent calls a save tool. The writing step sometimes adds things. A job-description comparison once came back with a market salary range, "$130k–$190k", that was in neither job description. A narrow check on that one workflow caught it; this is the general version.

I do have a second model that reads the output against the source, and it's useful. But its verdict is probabilistic, and if it decided whether a document gets saved, every one of its mistakes would become a failed run. So in my system anything that can end a run is deterministic, and the judge's verdicts go to an evaluation log that nothing in the run reads. That leaves a narrower question the code can answer: is this exact string in the source? For sentences that's useless, since every honest summary paraphrases. For "$4,200,000" or "§2.2" it's close to the right question. It won't notice a real figure attached to the wrong thing ("churn cost $4,200,000" when that was revenue); that stays the judge's job.

Where it sits in the agent loop

The check runs inside the save tool, before anything is written. It pulls claims out of the document by class, normalizes them (whitespace, case, curly quotes, thousands separators, "32 percent" against "32%") and looks for each one in the source text and in the fields the first model call extracted (the code calls those "anchors", and so does the message below). The extractors are deliberately stingy. A number only counts if it has four or more digits, a thousands separator, a decimal point or a scale like "10k", or if it's a section reference like "§12"; years from 1900 to 2099, list numbering, version strings and small counts never match.

If a claim isn't found, the save fails and the tool's result goes back to the model as a normal tool response. This is the message it gets, trimmed a little:

Artifact contains '$130k' — a currency claim that does not appear in this run's extracted anchors or source documents. Every specific figure and quotation must be grounded in the source. Fix it one of two ways: (1) replace it with the exact value from the source documents, or (2) if it is intentionally an unsourced estimate, keep it and append the literal marker [not in source] on the same line. Do not invent numbers, percentages, amounts, or quotations the sources do not state.

The model gets two ways out and has to pick one. After three rejections in a row the run ends. The marker is a bypass the model can use on anything, which I accept only because it stays in the saved document: an unsourced number can't be hidden, only labelled. The "Not specified" placeholder doesn't count as a marker, or "Not specified" next to an invented figure would sail through.

What it stopped on the first day

The real one came from the contract comparison. It was stopped for citing "§2.5" and "11.10" in contracts whose sections run 2.1–2.3 and 11.1–11.3. Those were real fabrications. The cause was upstream: the writing step had been given the extracted fields but not the contract text, so it was asked to cite clause numbers it had never seen (an earlier post is about that kind of bug). I gave it the text back. Writing the test for that fix turned up one more false positive: contracts number their headings "2.2 By Customer.", with no "§" and no "Section", so a correct "§2.2" would have been rejected too. Section claims now also match the bare number.

Another was a bug in my own regex. The financial summary wrote "Revenue reached $4,200,000, up 23%". My currency pattern was \d[\d,]*, which happily took the sentence's comma as part of the amount, so the claim became $4,200,000, and the save was rejected three times in a row, which ended the run, on a document that was entirely correct. Accepting commas only in groups of three digits (\d+(?:,\d{3})*) fixed it. The number extractor already did that; the currency one, written later, didn't.

The last one no regex could fix. With the comma fixed, the same summary stopped again, on $1,344,000. The source had revenue of $4,200,000 and a 32% gross margin, and this workflow is supposed to produce ratios and derived figures. $4,200,000 × 32% is its job. A verbatim check rejects every derived value by construction, and no pattern fixes that. That one decided the design.

Which claims it may block, per workflow

The $1,344,000 abort is why the check is configured per workflow rather than globally. Each workflow says which claim classes to check (section references are part of numeric) and what counts as a source. These are the two from that day, as they ended up; the financial summary is switched off with its old class list left in place:

contract_comparison:          {"enabled": true,  "claim_classes": ["numeric", "currency"], "sources": ["anchors", "extraction_source"]}
financial_statement_summary:  {"enabled": false, "claim_classes": ["currency"],            "sources": ["anchors", "extraction_source"]}

The contract comparison keeps numbers and amounts, because every figure in it should be copied from one of the two contracts. The financial summary doesn't run the check at all; its claims are still scored by the judge, and bad ones show up in evaluation instead of as stopped runs. As a rule of thumb:

Kind of claimWho checks it
Copied: amounts, large numbers, clause references in an extract or a comparisonThe verbatim check, and it can block
Derived: totals, ratios, conversionsNobody blocks; the judge watches
Paraphrased proseThe judge watches
Estimates the user asked forAllowed with the marker, visible as such

All of this replays offline, including the marker. The source is a two-line financial summary plus a contract heading, each case is one line the writing step might produce, and "first" and "fixed" are the patterns before and after that day (Python 3.11, standard library only):

--- first regex
prose comma       ['currency:$4,200,000,']
computed value    ['currency:$1,344,000']
computed, marked  pass
invented range    ['currency:$130', 'currency:$190']
real clause       ['section:§2.2']
invented clause   ['section:§2.5']
--- fixed regex
prose comma       pass
computed value    ['currency:$1,344,000']
computed, marked  pass
invented range    ['currency:$130k', 'currency:$190k']
real clause       pass
invented clause   ['section:§2.5']
The probe
import re

SOURCE = """Revenue: $4,200,000 (up 23% QoQ). Gross margin 32%.
2.2 By Customer. Either party may terminate on 30 days notice."""

FIRST    = re.compile(r"[$€£]\s?\d[\d,]*")
FIXED    = re.compile(r"[$€£]\s?\d+(?:,\d{3})*(?:\.\d+)?[kKmMbB]?")
SECTION  = re.compile(r"§\s?(\d+(?:\.\d+)*)")
MARKER   = "[not in source]"

def norm(s):
    return " ".join(s.split()).casefold()

def found(needle, corpus):
    return bool(needle.search(corpus)) if isinstance(needle, re.Pattern) else needle in corpus

def needles(kind, m, bare_section_needle):
    if kind == "section" and bare_section_needle:   # "§2.2" also matches "2.2 By Customer."
        bare = re.compile(rf"(?<![\d.]){re.escape(m.group(1))}(?![\d])")
        return (norm(m.group(0)), bare)
    return (norm(m.group(0)),)

def gate(artifact, patterns, bare_section_needle):
    corpus, rejected = norm(SOURCE), []
    for line in artifact.splitlines():
        if MARKER in line.casefold():
            continue
        for kind, rx in patterns:
            for m in rx.finditer(line):
                if not any(found(n, corpus) for n in needles(kind, m, bare_section_needle)):
                    rejected.append(f"{kind}:{m.group(0)}")
    return rejected or "pass"

CASES = {
    "prose comma":      "Revenue reached $4,200,000, up 23%.",
    "computed value":   "Gross profit was $1,344,000.",
    "computed, marked": "Gross profit was $1,344,000 [not in source].",
    "invented range":   "Market salary for this role is $130k-$190k.",
    "real clause":      "Termination is governed by §2.2.",
    "invented clause":  "Termination is governed by §2.5.",
}
for label, pats, bare in [("first", [("currency", FIRST), ("section", SECTION)], False),
                          ("fixed", [("currency", FIXED), ("section", SECTION)], True)]:
    print(f"--- {label} regex")
    for name, text in CASES.items():
        print(f"{name:17} {gate(text, pats, bare)}")

If your outputs are mostly derived, like forecasts or calculations, this check is the wrong tool. It's on in my development environment as of today, behind a switch, and has run on real documents for one day, so the numbers above are three runs, not a rate.