At the end of May, six of my document workflows kept inventing facts. A résumé said Acme Corp and the interview-prep output said TechCorp Inc. A job description said AlphaCorp and the comparison table said Not specified. I shipped two prompt changes, ran every workflow three times on byte-identical inputs, and the grounding score did not move.
The model had never seen the résumé. The step that was supposed to pull facts out of it was handed a tool summary that read 'parsed 6 sections', and nothing else. This post is for anyone about to rewrite a prompt for the third time. It makes three points, each with something you can reuse:
- Rerun the same inputs several times before and after a change. Without that you cannot tell a fix from noise.
- If the same wrong answer comes back on every run, the cause is deterministic. Check what the model received before you change what you ask it.
- Text gets lost in ordinary code between your data and the model. Here it happened in three places, and a short LangChain callback shows each one.
Three runs on identical inputs
Each workflow has an end-to-end acceptance test with several checks. One of them is a grounding score from 0 to 1: a separate grading pass compares the output with the input files and penalises anything not in them. Six workflows in the job-search category sat between 0.25 and 0.50.
The first fix was a universal prompt rule about using only facts from the source. The second was a set of per-workflow directives built on Anthropic's reduce-hallucinations guide: permission to say "not specified", quote-first extraction, a ban on outside knowledge. The pass rule was written into the plan before the first run: a score of at least 0.70 on two consecutive runs, because one good run can be luck.
The inputs were fixed files, verified byte-identical across runs with SHA256, so the only thing that could differ between runs was the model's sampling. Here is what came back:

Zero of six reached 0.70 on any run. Two workflows scored slightly lower after the second fix. And several fabrications were identical across runs. The section rewriter invented a project called "Real-Time Analytics Dashboard" in all three. The résumé analyzer said three times that the bullets had no numbers, when the résumé contained 70% and 50K events/sec. The comparison table wrote Not specified for both company names three times.
A repeated wrong answer points at something fixed
The commit that recorded these runs concluded that the model was falling back on its training prior: identical inputs, identical fabrications, so the model must prefer its own idea of a typical résumé to the one in front of it. That conclusion turned out to be wrong, and the reason is worth spelling out.
Anthropic's guide lists best-of-N as a check: run the same prompt several times, and if the outputs disagree, suspect hallucination. The converse does not hold: outputs that agree are not evidence that they are grounded.
The reasoning goes like this. Across the three runs, the inputs and the code were fixed and only the model's sampling varied. It really did vary: the runs used a temperature of 0.3, and other parts of the output changed from run to run (where the job description named the employer, one run wrote "Tech Industry" and the next "San Francisco Tech"). If the wrong answer does not change when sampling changes, sampling is not causing it. Something that stayed fixed is. That leaves two candidates: the model has a strong prior about this exact input, or the input it received is not the one you think it received. The second takes minutes to check by reading the request. I checked it later than I should have.
The structural fix, and why it still scored 0.30
With the prompt changes exhausted, the next step was the one the same guide recommends for long documents: extract word-for-word facts first, then write only from those. The composition step was split into two model calls. An extract pass pulls named anchors from the sources (company, dates, metrics) and a verifier checks each against the source text. A compose pass then writes the document from the anchors alone, with the raw tool outputs removed from its history. The design is in the prompt-chaining post.
The split ran end to end and worked as designed. The score was 0.30, the fourth sample at the same level, and the output now said TechCorp Inc. where the résumé says Acme Corp.
This time I read the extract pass's user message in the logs rather than its output. Its source blocks were these four strings:
'parsed 6 sections'
'extracted profile (skills={...}, _cache_hit=False)'
'extracted 177 words'
'fetch_knowledge failed: TOOL_INTERNAL_ERROR'
That was all. The extract pass collected its sources from the tool messages in the conversation, and the document tools returned short summaries there, not text. The verifier did its job: no anchor could be found in the sources, so every anchor became "Not specified". The compose pass received a block of "Not specified" values and a prompt that described a résumé it had never been shown, and it produced a plausible one.
The three earlier runs had no extract pass, but the agent worked from the same tool messages, so it very likely had no more to go on. I can't show that, because those requests were never logged. What I can say is that the "training prior" explanation was never tested against the simpler one: the model was filling in a document it had been told about and never given.
Giving it the text
Fixing the tool summaries alone would not have been enough. In the failing run, a rule in my agent that allows only certain tools at each stage of the workflow had blocked one of the document parsers, so whether the raw text shows up in any tool message depends on what the agent happened to call. The fix therefore does not depend on the agent's tool calls. Every job now parses its input files to text before the agent loop starts and keeps the result in graph state, and the extract pass reads its sources from there, with the raw file text placed before any tool summary so the first thing the model reads is the source itself. Each file's text is capped at 30,000 characters, with a marker where it was cut, so the state and its checkpoints stay small. The rerun, again on the same inputs:
| Workflow | Before (3 runs) | After (1 run) |
|---|---|---|
interview_prep | 0.35 / 0.35 / 0.30 | 0.72 |
resume_analyzer | 0.40 / 0.40 / 0.40 | 0.81 |
resume_section_rewriter | 0.30 / 0.30 / 0.30 | 0.72 |
jd_comparison_table | 0.30 / 0.25 / 0.25 | 0.70 |
The artifact now said Acme Corp. Two caveats, stated plainly. Each "after" is one run, so these are not the two consecutive passes my own rule asked for. And it did not hold everywhere.
The next round used the same method one step further along: log each step's input and output, and find the first step where the wrong fact appears. With Bedrock request and response logging switched on, the extract pass's anchors were correct, so the wrong facts first appeared in the compose pass. It added an unanchored salary band and turned a 70% deploy-time figure into 70% test coverage. A prompt-only fix for the composer was measured and failed as well. What worked was changing what the composer received and what it was allowed to output: more specific anchors, and patterns the output validator rejects. That flipped three workflows to passing. Even so, by June 4 the section rewriter was back at 0.40 with the same fabrication signature on three samples, and it and the interview-prep workflow were taken out of the regular test runs again. Raw text was necessary. It was not sufficient.
The same bug in two other places
Once I was looking for it, the same class of bug turned up twice more in one day, June 5, on a product-page workflow.
The first was in the agent's tool node. My MCP tools return a short summary and the actual output under data, and the code that built the ToolMessage forwarded only the summary. The parser's output reached the model as the string "parsed 2793 chars / 6 words", so the product URLs in the file never entered the conversation. The agent behaved accordingly: it said it was unable to read the file, asked to read stored results by reference ids it had made up, and parsed the same file again. The fix appends the data payload to the message.
The second never reached the model at all. A new version of the system prompt included a JSON example, and the prompt is rendered with Python's str.format. The loader caught the resulting exception and fell back to a generic prompt. Here is the shape, reduced to a script you can run:
SYSTEM_PROMPT = """You extract product data from {file_count} uploaded file(s).
To read a file, call parse_document with {"file_id": "<the uploaded file id>"}.
Never call read_scratchpad."""
def render(template, **variables):
try:
return template.format(**variables)
except Exception as e:
log.warning("Failed to load prompt, using fallback: %s", e)
return GENERIC_FALLBACK
WARNING Failed to load prompt, using fallback: '"file_id"'
'You are a helpful assistant. Complete the workflow.'
str.format reads {"file_id" as a field name and raises KeyError; the dev worker logged the same line word for word. None of the new rules in that prompt version reached the run, and the agent did exactly what they forbade. The fix doubled the braces. The guard is a test that renders every system prompt with the runtime variables, because the existing test only checked that each prompt loaded, and loading succeeded.
In all three cases (the extract pass's sources, the tool message, the prompt) the model's behaviour was a reasonable response to what it was given. What it was given was wrong, and nothing between the code and the model said so.
Printing what the model received
The cheapest tool here is a callback that prints every message at the moment it is handed to the chat model. LangChain calls on_chat_model_start with the final message list, after your prompt templates, trimming and message assembly have run. This runs offline with langchain-core 1.4.0, the version in production at the time, and a fake chat model:
class WhatTheModelSaw(BaseCallbackHandler):
"""Print every message exactly as it is handed to the chat model."""
def on_chat_model_start(self, serialized, messages, **kwargs):
for batch in messages:
for m in batch:
text = m.content if isinstance(m.content, str) else json.dumps(m.content)
print(f"{m.type:>6} {len(text):>5} chars | {text[:70]!r}")
tool_result = {"summary": "parsed 2793 chars / 6 words",
"data": {"extracted_text": "https://shop.example/p/1
https://shop.example/p/2
..."}}
def history(tool_content):
return [SystemMessage("Extract product metadata from every URL in the uploaded file."),
HumanMessage("Here is my file."),
AIMessage("", tool_calls=[{"name": "parse_document", "args": {"file_id": "f1"}, "id": "t1"}]),
ToolMessage(tool_content, tool_call_id="t1")]
llm = FakeMessagesListChatModel(responses=[AIMessage("ok")] * 2)
llm.invoke(history(tool_result["summary"]), config={"callbacks": [WhatTheModelSaw()]})
llm.invoke(history(tool_result["summary"] + "
" + json.dumps(tool_result["data"])),
config={"callbacks": [WhatTheModelSaw()]})
-- ToolMessage built from result['summary'] only
system 61 chars | 'Extract product metadata from every URL in the uploaded file.'
human 16 chars | 'Here is my file.'
ai 0 chars | ''
tool 27 chars | 'parsed 2793 chars / 6 words'
-- ToolMessage built from summary + data
tool 105 chars | 'parsed 2793 chars / 6 words\n{"extracted_text": "https://shop.example/p'
A 27-character tool message after a document parse is the whole diagnosis in one line. In a real graph, attach the handler through the config you already pass to ainvoke or astream.
The callback shows LangChain messages, which answers "did my code build the right messages?". It does not show the provider request, which LangChain builds afterwards from those messages. To see that on Bedrock, set the logger langchain_aws.chat_models.bedrock_converse to DEBUG; it logs each request and response. My logging setup silences the langchain loggers in production, so one environment variable raises only that logger, on dev only, because the requests contain user documents.
What to do before the next prompt rewrite
To diagnose, in this order:
- Fix the inputs and hash them. Decide the number of runs and the pass rule before you look at any result.
- Run at least three times. If the scores and the wrong answers vary, you are looking at sampling noise: add runs, not prompt text. If the same wrong answer repeats, stop editing the prompt; the cause is something that stayed fixed.
- Pick one fact you know is in the source, such as a company name, and search for it in the request the model received (the callback above, or the provider's DEBUG log). If it is not there, no instruction can make the model use it.
- If it is there, log the input and output of each step and find the first step where the wrong fact appears. Fix that step, not the last one.
Guards worth adding to any agent codebase, one per place text was lost here:
- Tool results: put the tool's data in
ToolMessage.content, not only its summary. If the data is too large, store it and put a reference the model can actually read in the message. A test can assert that a known string from the tool output appears in the message. - Prompt rendering: render every prompt template with the real runtime variables in a test. A render error should fail the run, not switch to a fallback prompt without anyone noticing.
- Source collection: when one step gathers sources for another, give it the original documents, and put them before any summary of them.
This does not help when the outputs vary from run to run. That is a sampling problem, and it needs more runs and a tighter pass rule, not a request dump. And when the fact is genuinely missing from the source, the right output is an explicit "not specified", which is a separate problem from the one here.

