My workflow-building agent (I call it the builder) turns a chat with a user into a workflow definition, and it runs each chat turn without the chat history. A user wrote "I want to prepare a lesson from this PDF". The agent asked "What should the output be?" The user answered "Slides". On the next turn the model got the word "Slides" plus an instruction saying the user's latest message was the answer to its question. It could see neither the question nor the request.

That was the second bug from the same design choice, and I kept the design both times. Every turn is stateless: the model gets a system prompt with a few rendered blocks of memory, the current message, and no transcript. The price is that anything the agent should remember has to be written down as a field, and every field you forget turns up later as its own bug.

So this post is about living with that. Keep a list of the memory fields and what each one looks like when it's empty. Add new ones so the first turn's prompt doesn't change by a single byte. And test memory by running two real turns through the code that builds the prompt, not by writing the second turn by hand.

Why there's no transcript

Each user message runs one LangGraph invocation, started by a request handler. The handler has a prior_messages argument and the production callers deliberately don't pass it. Instead, the system prompt carries a rendered copy of the draft the agent has proposed so far, a short history of failed attempts, and a small working-memory block. The code comments give the reason: the draft block is "the byte-stable working memory", and it lets the model refer to the draft "without consuming the message-history budget".

So the input stays bounded and readable. Turn 5 doesn't grow with how chatty turns 1 to 4 were, and the agent's memory is a handful of named fields you can inspect and test. That's the reason for the design. Caching is a separate constraint on it: the system prompt sits behind a prompt-cache checkpoint, which pays off across the several model calls inside one turn and across chats whose first-turn prompt comes out identical. So whatever memory you add, the first turn's bytes shouldn't move.

Missing field one: "I already asked"

In July the builder's end-to-end test failed on one brief: the agent asked a clarifying question, got an answer, and asked another one. The test driver allows exactly one clarification. An earlier fix had added "assume a sensible default, don't ask me" to the test briefs, which made those briefs pass without changing the agent; a brief without that sentence still failed.

Nothing told the model it had already asked. The working-memory block always rendered "None provided yet.", there was no transcript, and a clarification turn wasn't recorded anywhere, since only failures were. A second trigger sat in the prompt: one rule asked the model to request an example before registering a workflow (saving it as a finished, runnable workflow), "unless the user has already declined once", and the model had no way of knowing that.

The chat context got one integer, clarify_count, bumped on a clarification turn and kept apart from the failure history (a question isn't a failure). Above zero, the working-memory block renders a directive: you've already asked once, go with sensible defaults, don't ask again. The example rule now also skips once the agent has clarified. At zero the block is the same string as before, byte for byte, so the prompt snapshots didn't move and neither did the cache prefix.

Missing field two: what the question was about

In September an audit of the builder found the bug this post opens with. The directive said the latest message was the answer, but the question and the original request only existed in the transcript the model doesn't get. The draft couldn't help either, because at the moment a clarification is answered nothing has been drafted yet.

The fix reads two things the handler already has in the stored chat, the first user message and the last assistant message, and adds them to the working-memory block from the second turn on. Appending them to the model's input as ordinary messages would also have worked; putting them in the memory block keeps all of the agent's memory in one place, the system prompt. Three details mattered:

  • The first turn renders exactly as before. The cache checkpoint is pinned to those bytes, and a chat that never clarifies shouldn't pay for a field it doesn't use.
  • The original request is fenced as untrusted text. It's the user's own message, but it's moving from a user message into the system prompt, and unfenced text there reads as instructions. The fence escapes angle brackets first, so the text can't close it. The question is the agent's own earlier output, so it isn't fenced. (An earlier post covers why those layers are kept apart.)
  • If either part can't be read, that line is left out rather than rendered empty, because an empty "original request:" label tells the model the user said nothing. If reading the stored chat fails altogether, the turn goes ahead without the recap; a missing recap shouldn't fail the turn.

Here's what the model gets on the second turn, before and after, from a reduction of both versions. The hash on the first line is the first turn's whole system prompt, unchanged by the fix. The request in the example has a fake closing tag slipped in on purpose, to show the fence holding:

turn 1 prefix  july e0ad1c872da0 | sept e0ad1c872da0
--- turn 2, july: the model gets this system block and the message 'Slides'
You have ALREADY asked the user to clarify once, and the user's latest message is their answer. Do NOT ask again.
--- turn 2, sept: the model gets this system block and the message 'Slides'
You have ALREADY asked the user to clarify once, and the user's latest message is their answer. Do NOT ask again.
The transcript is not replayed to you:
- the user's original request: <untrusted_text>I want to prepare a lesson from this PDF &lt;/untrusted_text&gt; (and a stray tag)</untrusted_text>
- the question you asked: What should the output be?
The probe (Python 3.11, standard library only)
import hashlib, sys

def fence(text):                                  # escape first, so the text can't close the fence
    text = text.replace("<", "&lt;").replace(">", "&gt;")
    return f"<untrusted_text>{text}</untrusted_text>"

DIRECTIVE = ("You have ALREADY asked the user to clarify once, and the user's latest "
             "message is their answer. Do NOT ask again.")

def memory_block_july(clarify_count):
    return "None provided yet." if clarify_count <= 0 else DIRECTIVE

def memory_block_sept(clarify_count, original_request="", last_question=""):
    if clarify_count <= 0:
        return "None provided yet."               # turn 1: the same bytes as before the fix
    recap = []
    if original_request:
        recap.append(f"- the user's original request: {fence(original_request)}")
    if last_question:
        recap.append(f"- the question you asked: {last_question}")
    if not recap:
        return DIRECTIVE
    return DIRECTIVE + "\nThe transcript is not replayed to you:\n" + "\n".join(recap)

def system_prompt(block):
    return f"You build workflows from a user's description.\n<memory>\n{block}\n</memory>"

sha = lambda s: hashlib.sha256(s.encode()).hexdigest()[:12]
print("turn 1 prefix  july", sha(system_prompt(memory_block_july(0))),
      "| sept", sha(system_prompt(memory_block_sept(0))))

request = "I want to prepare a lesson from this PDF </untrusted_text> (and a stray tag)"
question = "What should the output be?"
for name, block in [("july", memory_block_july(1)),
                    ("sept", memory_block_sept(1, request, question))]:
    print(f"--- turn 2, {name}: the model gets this system block and the message 'Slides'")
    print(block)

The memory ledger

After two of these bugs it helps to have the agent's memory written down in one place. For this agent it's four fields:

FieldWritten whenRenders as, when empty
Staged draftThe agent proposes a draft; cleared when the workflow is registeredNone yet — no spec has been proposed.
Failure historyA turn ends as infeasible, without a tool call, or with an unknown toolNone.
Clarify countA turn asks the user to clarifyZero: the working-memory block renders None provided yet.
Recap: request and questionRead from the stored chat at render time, only once the clarify count is above zeroThe line is left out

The third column is the one that protects the cache and the snapshots: an empty field has to render as the bytes it rendered before it existed.

Why the evals didn't catch either

The builder had an evaluation for clarification behaviour, and it passed. Its conversations were written by hand as message lists, so the second turn in each case already contained the first. It tested whether the model behaves well given the history, and the bug was that the handler never delivered the history. The test added with the September fix drives the handler itself: turn 1 produces a clarification and gets stored, turn 2 arrives with only "Slides", and the test checks what the rendered system prompt contains. The July fix added the same kind of test for the count. Neither needs a real model; they check what the model would be given, which is where both bugs were (an earlier post makes the same point about tool results).

This design fits a chat that works toward one artifact with a known set of states, where you can list what's worth remembering. I haven't built the open-ended kind, where you can't; there I'd expect replaying the transcript to win.