In May, four of my agent workflows each ran for between three and seven minutes and called no tools at all. None of them raised an exception, and all of them stopped short of the 840-second timeout. In July another agent read the same stored result 45 times, with a different search term each time, until the per-run spending cap stopped it. The only signal any of these runs gave was what they cost.

This post is for anyone running a LangGraph agent loop in production. It makes three points:

  1. A stuck agent does not fail. Exceptions, timeouts and iteration caps either never fire or fire only after most of the budget is gone. You need detectors that look for missing progress.
  2. Each kind of loop varies something. A detector that compares on the thing the loop varies never fires. Count something the loop cannot vary.
  3. Keep one hard budget that stops the next turn as the last line of defence, and stop the graph cleanly when any detector fires.

Below are five loops from my own runs, what each one varied, and what finally caught it. Three short scripts reproduce the detector behaviour offline with langgraph 1.2.7 and langchain-core 1.4.8, the versions in production at the time.

Why the usual limits miss it

A loop produces normal-looking turns: the model answers, a node runs, the graph moves to the next step. Nothing raises. The limits you would expect to catch it each have a gap:

  • The wall-clock timeout was 840 seconds, the Lambda ceiling minus a margin. A loop that uses 400 seconds is well inside it.
  • The iteration cap and LangGraph's recursion_limit count steps. My runner sets recursion_limit to three times the workflow's turn cap plus ten. With model turns that take ten seconds or more, the cap arrives only after minutes of spending.
  • The stuck-state detector aborted a run when the same tool was called five times in a row with identical arguments. That catches an agent repeating itself exactly. Every loop below varied something.

Loop 1: talking instead of acting

The four May runs cycled between the reasoning node and a nudge that asked the model to call a tool. The model answered in prose each time and never issued a tool call. From outside, the run was busy. None of them ended in an error; the records I kept don't say which limit finally ended each one.

The fix is a watchdog on the stream. After each super-step (one round of node executions), count the tool calls the model has requested plus the tool results that have come back. If that count does not grow for 60 seconds, stop the run. Reduced to a runnable script:

from langgraph.errors import GraphDrained
from langgraph.runtime import RunControl          # both in langgraph 1.2

def count_dispatches(messages) -> int:
    """Tool calls the model asked for, plus tool results that came back."""
    n = 0
    for m in messages:
        if isinstance(m, AIMessage):
            n += len(m.tool_calls or [])
        elif isinstance(m, ToolMessage):
            n += 1
    return n

control = RunControl()
last_count, last_progress = None, time.monotonic()
try:
    async for state in graph.astream(inputs, config, stream_mode="values", control=control):
        count = count_dispatches(state["messages"])
        now = time.monotonic()
        if last_count is None or count > last_count:
            last_count, last_progress = count, now
        elif now - last_progress > idle_s and not control.drain_requested:
            control.request_drain("no_progress")
except GraphDrained:
    snap = await graph.aget_state(config)

With a fake reasoning node that takes 0.2 seconds per turn and a 1-second threshold:

watchdog: no tool dispatch for 1 s after 5 model turns
drained at 1.0 s, reason=no_progress, checkpoint next=('reason',)

Two details matter. First, the count starts from the first chunk, not from zero, so a resumed run with old tool calls in its history is not treated as progress or as a stall. Second, how the run stops. The script does not kill the task; RunControl.request_drain asks LangGraph to stop at the next super-step boundary, which raises GraphDrained and leaves a checkpoint that can be resumed or inspected. GraphDrained carries no reason of its own, so read it back from control.drain_reason. My runner uses the same drain for a user's Stop button. Stopping at 60 seconds instead of at the 840-second timeout was estimated to cut model spend on runs like these by about 13 times; that is the ratio of the two limits, not a measurement.

One limit of this watchdog: it checks when a chunk arrives from the stream. A single node that hangs produces no chunk, so it never trips; the wall-clock timeout still has to cover that case.

LangGraph 1.2 also has TimeoutPolicy(idle_timeout=…) on add_node, and it looks like the built-in version of this. It isn't, and the reason is worth knowing. The idle timer belongs to one attempt of one node and restarts with every attempt. This loop is made of many short attempts across super-steps, so no single attempt is ever idle for long. The same two-node loop with a 1-second idle timeout on both nodes:

policy = TimeoutPolicy(idle_timeout=1.0)
g.add_node("reason", reason, timeout=policy)   # 0.2 s per turn
g.add_node("nudge", nudge, timeout=policy)
stopped by GraphRecursionError after 5.4 s

The idle timeout never fired; only the step limit ended the run. And if the node has a RetryPolicy with the default retry_on, a NodeTimeoutError counts as retryable, so a timeout that did fire would run the stuck node again. I kept the watchdog.

Loop 2: the same call with different arguments

The July run was an expense workflow. A tool had stored its large result and returned a reference, and the agent was supposed to read it back with read_scratchpad. Each read returned nothing, because the code that builds the tool message looked for the payload under data while this tool returned it under content. That is the same class of bug as in the previous post. The agent did what a reasonable reader would: it tried again with a different search term or a different length. Because the arguments changed every time, the identical-call detector never fired, and because every step dispatched a tool, the watchdog saw progress.

The first fix surfaced the payload and added a cap of 20 reads per stored reference. The rerun still looped, 45 reads, until the spending cap. Two things were wrong. The workflow's own instructions told the agent to search the stored result by vendor name, and a stored result has no pagination, so each search returned a partial slice and the agent kept searching. And the per-reference cap counted only reads that named a reference. These reads didn't, so the cap never counted them.

The second fix told the agent to read the whole result in one call, and added a cap on the total number of reads that ignores the arguments completely. This script replays the loop, 45 reads that differ only in their search term and length, against four detectors:

  • same_call_streak: stop when one tool is called five times in a row with identical arguments.
  • dispatch_watchdog: stop when the number of tool calls stops growing.
  • per_ref_cap: stop after 20 reads of the same named reference.
  • total_cap: stop after 20 reads, whatever the arguments.
calls = [("read_scratchpad", {"ref_id": "", "grep": vendors[i % 5], "max_chars": 4000 + i})
         for i in range(45)]
same_call_streak   never fires in 45 calls
dispatch_watchdog  never fires in 45 calls
per_ref_cap        never fires in 45 calls
total_cap          stops the loop at call 21
A grid of 45 squares, one per read_scratchpad call. Three example calls above it differ only in their grep and max_chars arguments. Squares 1 to 20 are indigo, square 21 is solid red and labelled total read cap: stops at call 21, and squares 22 to 45 are faded, ending at a label reading without it: spending cap at call 45. A line under the grid says the identical-call check never fires because the arguments always differ.
The read loop, replayed. The example arguments come from the reproduction script, not from the production run.

The general rule: a detector keyed on the call's arguments can be defeated by any change in the arguments, and a model that is failing will change them. Count per tool name, independent of the arguments, and set the cap well above what a legitimate run needs.

Loop 3: searching instead of working

A product-page workflow ran 30 search_tools calls, parsed its input twice, and never once called the tool that fetches a product page. The spending cap stopped it. The workflow is divided into phases, each meant to list the tools allowed in it, and this workflow's phases listed none, so on every turn the agent had nothing bound that pointed at the next step and went back to searching the whole catalogue. Each search used a different query, so this loop too varied its arguments and dispatched a tool on every step. Declaring the allowed tools for each phase fixed it. The tool search post covers the cost side of letting the model discover its own tools.

Loop 4: a gate the agent cannot satisfy

My agent has a phase state machine: each phase allows certain tools, and a middleware rejects a call to a tool the current phase doesn't allow. One run made 36 rejected calls and produced nothing. Two bugs combined to cause it.

The agent works out its current phase from the last tool it called. Two things corrupted that input:

  • The tool node recorded "last tool called" before the middleware decided whether to allow the call. A rejected call therefore still moved the phase: in one run it went from research back to extract, then jumped ahead to compose.
  • Tools that belong to no phase, such as search_tools, overwrote the same record. The lookup then fell back to the first phase after every search.

On top of that, the list of tools offered to the model and the check that rejects calls each computed the phase separately, so they could disagree. The model was offered tools that the check then rejected.

Every rejection comes back to the model as a tool result, so a dispatch-count watchdog sees each one as progress. What fixed it: write state only after the check passes, give the phase its own field that only real domain calls update, and derive both binding and dispatch from one function.

Loop 5: output cut off at max_tokens

Some education workflows send a whole document as the arguments of a single tool call. In one run, turns 3 to 8 each produced exactly 8,192 output tokens, the configured maximum, and took about 115 seconds each. The tool call at the end of each turn was cut off, arrived with empty arguments, failed validation with Missing required keyword only argument, and the model tried again, until the 840-second timeout. The first reading of the failure was that the model had formatted its call carelessly, and a prompt change was proposed for that.

The signal was in the response the whole time. Bedrock's Converse API reports why the model stopped, and LangChain passes it through as response_metadata["stopReason"]; a value of "max_tokens" means the output was cut off (the Converse API reference lists the values). Anthropic's own API reports the same as stop_reason. A tool call from a truncated turn is not a careless call. It is an incomplete one, and asking the model to try again produces the same truncation. My fix raised the cap to 16,384 and kept the existing warning when output comes close to it; a circuit breaker that stops retrying truncated calls was recorded as a follow-up.

The spending cap is the backstop

LoopWhat it variedWhat caught it
Talking instead of actingNothing; it made no tool callsNothing at the time; the watchdog came after
Same call, different argumentsThe argumentsSpending cap, then a total read cap
Searching instead of workingThe search querySpending cap
A gate it cannot satisfyThe tool it triedNothing; the run failed with no output
Output cut at max_tokensNothing; each turn hit the limitThe 840-second timeout

Two of the five were stopped by the per-run spending cap and nothing else. The cap works in a simple way: the turn that crosses the limit finishes and is billed, and the next turn is refused before any model call. It is deterministic code, not a model judgement, and it cannot be evaded, because every turn of every loop costs money. Its weakness is that it fires only after the money is spent. It is the last line of defence, not detection.

What to add to your own loop

  1. A no-progress watchdog on the stream: count tool calls and tool results per super-step, and stop when the count has not grown for longer than your slowest legitimate turn. Measure that turn first; a 60-second threshold is too short if a single composition turn can take two minutes.
  2. Per-tool caps that ignore the arguments, for any tool that can be called repeatedly (reads, searches, retries). Keep the identical-call detector as well; it catches a different, cheaper failure.
  3. A check on stopReason (or stop_reason) after every model turn. Treat a tool call from a max_tokens turn as invalid instead of passing its error back to the model.
  4. A rule for gates: state that a check reads is written only after the check passes, and the tools the model is offered and the tools the check allows come from the same function.
  5. For loops 4 and 5, which look like progress to the watchdog: a cap on consecutive rejected or failed tool calls. I haven't built this one, so treat it as a suggestion rather than a tested rule.
  6. A per-run spending cap that refuses the next turn, as the backstop.
  7. When any of these fires, stop with RunControl.request_drain(reason) rather than cancelling the task, and record the reason as the run's terminal state, so a stopped run can be told apart from a failed one.

None of this replaces fixing the cause. Loops 2 and 3 started with the agent being given the wrong thing: an empty tool result, instructions that named the wrong strategy, no tools for the phase it was in. Loop 4 was a bug in my own gate, and loop 5 a limit set too low for the job. The detectors limit the damage while you find that cause.