My workflow-building agent has a tool that assembles a prompt from named blocks in a library. In a test conversation last week, its first call asked for blocks called role_intro, task_instruction and output_format. None exist. Its second call tried a category, document, that doesn't exist either, and after that the turn was out of tool calls, with no result. The tool's rejection messages did list the valid values, but only after a call had been spent. Its description, the only thing the model reads before the first call, just said the id must be "registered in the block library". So the model made up plausible names.

That was the fifth place in a single week where the tool names my agent worked from came from text I'd written instead of from the registry of tools that actually exist. Hand-written names are copies with no link back to the original, so they drift, and an agent reads them as instructions. Once I'd found all five, they split into two kinds with two different fixes. Names the agent reads at runtime get generated from the registry. Names that only tests read get checked against it. The price is longer tool descriptions on every call, in exchange for not paying a model round per guess.

Where the name was writtenWho reads itFix
1. The inventory tool's listThe agent, at runtimeGenerated from the catalog
2. The system prompt's example idThe agent, at runtimeCorrected
3. Tool descriptions (block tool, template tool)The agent, at runtimeGenerated from what's on disk
4. Eval goldens and their harnessTests onlyScanned against the catalog
5. Two retrieval corporaRetrieval, no one reviews itReplaced with real names

Names the agent reads: generate them

Three of the five were things the model sees while it works.

The inventory tool. The agent's prompt shows a shortlist of tools, grouped by the MCP server they live on, and servers that don't make the shortlist are folded into one line with an instruction to call an inventory tool for the rest. That tool returned a hand-written list of 24 names, and 20 weren't in the catalog. So the recovery path, the thing the agent is told to use exactly when the shortlist isn't enough, handed out fake names. Following it ended at registration, the check that runs when the agent saves a finished workflow and rejects any tool the catalog doesn't have.

The prompt's own example. The system prompt tells the agent to use "the exact canonical id (e.g. 'resume-profile-mcp:parse_document')" and warns that made-up ids get rejected. That example didn't exist; the real tool is parse_resume_document. I have no evidence the agent ever copied it, but it was the one id the prompt offered as a correct answer.

Tool descriptions. The block tool above, and a template tool whose argument was described by a naming pattern and five examples ("such as 'document__txt', 'video__srt', ...") and whose rejection listed nothing.

The fixes all go the same way. The inventory tool no longer has a list; it reads the catalog through the same function that renders the prompt's shortlist and that registration uses, so the three can't disagree. The hand-written list was deleted, not kept as a fallback. The prompt example now names the real tool. And both tools' descriptions are built from what's on disk, as is the template tool's rejection:

_TEMPLATE_IDS = ", ".join(repr(t) for t in list_available_templates())

description=f"'<input>__<output>' key, one of: {_TEMPLATE_IDS}."
...
return _refusal(f"no phase template {template_id!r}; available: {_TEMPLATE_IDS}")

The block tool got the same: its description lists every block id and category in the library. Both places need the list, for different calls: the description is all the model has read before its first call, and the rejection only helps the second. The next test turn after the block tool's fix made its calls without a guess and ended with a draft. Longer descriptions cost tokens on every call, softened here because the tool list sits behind a prompt-cache checkpoint.

Names only tests read: check them

The other two were text the model never sees in production, which is why nothing noticed.

The eval corpus. The agent's golden test conversations and their harness named 47 distinct tool ids, and 34 weren't in the catalog. Three of those were supposed to be missing: one golden checks that the agent can report tools it would need but doesn't have. That leaves 31 invented ids, like pdf-mcp:to_docx and email-mcp:send, spread over 64 places in the files. The eval replaces the model with canned replies, and those replies never reach registration, which would have rejected them.

Two small retrieval corpora, one on the tool server and one in the agent, had tool names inside their entries. On the server, 12 of the 14 were made up, and the agent's copy had the same problem. Both now use real tool names, one per tool.

A golden conversation has to spell out tool ids, that's what it is, so it can't be generated. Instead the eval directory has a test that scans every .py, .yaml and .json file for anything shaped like a tool id and checks it against the catalog. It scans text rather than loading each golden, so a new golden file is covered without anyone registering it. Three details decide whether it can actually fail. It has a floor, so a pattern that finds nothing is a failure, not a pass (an earlier post is about that trap). Exemptions are derived from the golden that declares them, and a second test asserts they really are absent from the catalog, so the exemption list can't quietly cover a real tool. And the reference is independent: the inventory's test compares against the raw catalog, not against the function the inventory now calls, because comparing a function with itself proves nothing.

A small version of the scan, on one prompt line and one golden line (Python 3.11, standard library). The first two output lines show the floor; the third finds both fake ids:

import re

REGISTRY = {"pdf-mcp:extract_text", "resume-profile-mcp:parse_resume_document",
            "doc-render-mcp:to_pdf", "email-draft-mcp:compose"}
PROMPT = "Use the exact canonical id (e.g. 'resume-profile-mcp:parse_document')."
GOLDEN = "steps: [pdf-mcp:extract_text, pdf-mcp:to_docx, email-mcp:send]"

def check(texts, pattern, floor):
    found = set().union(*(set(re.findall(pattern, t)) for t in texts))
    if len(found) < floor:
        return f"FAIL: scanned only {len(found)} ids, expected at least {floor}"
    bad = sorted(found - REGISTRY)
    return f"FAIL: not in registry: {bad}" if bad else f"ok ({len(found)} ids)"

WRONG = r"\b[a-z_]+_mcp:[a-z_]+\b"                  # assumed "pdf_mcp:...", finds nothing
FULL = r"\b[a-z0-9]+(?:-[a-z0-9]+)*-mcp:[a-z0-9_]+\b"
print("wrong regex, no floor:", check([PROMPT, GOLDEN], WRONG, floor=0))
print("wrong regex, floor 3: ", check([PROMPT, GOLDEN], WRONG, floor=3))
print("right regex, floor 3: ", check([PROMPT, GOLDEN], FULL, floor=3))
wrong regex, no floor: ok (0 ids)
wrong regex, floor 3:  FAIL: scanned only 0 ids, expected at least 3
right regex, floor 3:  FAIL: not in registry: ['email-mcp:send', 'pdf-mcp:to_docx', 'resume-profile-mcp:parse_document']

Before the fix the real scan failed in those 64 places. The goldens were rewritten with real ids, and steps the catalog can't do at all, like sending email or posting to Slack, became steps it can, keeping each golden's shape (a tool added, removed or swapped) so it still tests the same behaviour.

None of this stops a model from inventing a name nobody wrote; registration catches that, and it worked the whole time. The names I wrote myself were the ones steering the agent toward it, and those are fixed now.