Our builder is an agent: you describe a task ("list the deadlines in every contract I upload") and it gives back a workflow you can run again on new files. The harness is everything we wrap around the model: the prompt, the tools, the checks on what it proposes, and the tests.

This month we found two loops in that harness, and both ate tokens. We're still building our fix, so what follows is directions, not code.

Loop one: a conversation that never ends

On September 29, in a dev test, someone asked the builder for a daily news summary. Five turns later, nothing was saved:

TurnUser saidModel calls
1Summarise the news on a topic1
2Finance (answering "which topic?")5
3Looks fine4
4Skip the examples5
5Looks fine3

The model showed a draft, the user said fine, it asked for example files, the user said skip, and it showed the same draft again. 18 calls. The prompt said "Do NOT call clarify_requirement again". The model asked again.

Why it went in circles:

  • The model saw only the latest message. That was on purpose, so everything else had to come from saved state.
  • The saved draft kept only names and tool lists. To confirm it, the model had to write the whole draft again.
  • Nothing saved "waiting for OK" or "user said skip".
  • "Ask only once" was just a sentence in the prompt.

Smaller versions were everywhere. A tool description didn't list its valid ids, so the model guessed twice (told here). The normal path already took six tool calls, so two misses pushed the turn past its limit of seven. The user got "unexpected response". Another tool refused to save a file because the model left the file type at its default. The refusal was one line, so the model sent the same call again until the run was stopped. That happened in all three runs. The pull request that fixed it says: "The model never reads a one-line refusal as 'set this argument'."

What to do

  • Keep "where are we", "what did we already ask" and "what did the user say no to" in state that your code checks.
  • When a tool rejects a call, say the exact field and the value it needs. If only one answer makes sense, like a default the model left out, fix it yourself and go on.
  • Set a call budget per turn and per chat, leave room for at least one retry, and count calls per finished result.

Watch out

  • A shorter prompt is not the big win it looks like. Every call re-sent about 20,700 tokens, but about 80% were read from the prompt cache at around a tenth of the price. We cut the prompt text by 18% and the tool descriptions by 57%, and that cut the cost per call by only about 3% and 6%. The number of calls is the cost.
  • Before you blame the model, check the tool. One "the model guessed wrong" was really our tool's container shipping without its data folder.

Loop two: every wall adds a layer

Looking back at four months of builder work: we had no real users, so goals came from our design docs. Our own tests passed, the feature stayed off in production, and when a rare real run hit a wall, we added one more layer. Then it started again.

A closed loop of six steps: no real users, goals from design docs, our own tests pass, feature off in production, a real run hits a wall, and in red, one more layer. In the middle: back to the top, a little bigger.
Four months of builder work, drawn as one loop.

What told us it was a loop:

  • The first real end-to-end run needed six fixes in one day. Unit tests missed all six. One was a missing permission that mocks don't check, one was an argument the test skipped, one test checked the wrong idea.
  • The same bugs came back: lost chat context three times, duplicate saves three times, and four times we found the self-test (the builder running a new workflow once to check it) had never really run.
  • We once used "how many tools a step uses" to decide if a workflow could run as fixed steps, with no LLM. It passed every test we wrote. It was wrong, and a person caught it.
  • In production, no one has run a workflow the builder made yet.

And loop two feeds loop one. Every layer added tools and checks, and each one is another call the model can fail. That's how the normal path grew to six tool calls before a draft.

What to do

  • Before you optimize, pick real requests, real files, and a person who judges "can I use this".
  • Every tool output should be used by a later step or by your code. If nothing uses it, drop the call. In one good turn, two of six calls made output nobody read.
  • Give the work an end date and a default, like "if real users don't show up by then, we stop". And let someone else make that call. In a lab study (Boulding et al., 1997), managers who wrote their own stop rules stopped a failing product 0% of the time.

Watch out

  • Tests you write mostly prove the code matches your own description, not that it's useful to anyone.
  • A small test set can't see small wins. By our simulation, 15 tasks run three times can't reliably see a pass-rate change under about 19 points. So a prompt tweak that "looks better" there is not progress yet.

Where we are

We've frozen new builder features, with an end date and a default. First we check if real users repeat the same task often enough to want a saved workflow. If not, the builder becomes an internal tool for building our own workflows. We don't have the answer yet.