Put a paid API call and an interrupt() in the same LangGraph node, and resuming it sends the paid call again. I ran it: a human approved gpu-job-1, the node re-ran from its first line on resume, submitted gpu-job-2, and the final state recorded gpu-job-2 as approved. Nothing raised. My own agent had already hit two relatives of it. In late August, a flag turned one human approval into approval for every later high-risk call in the run. In June, a Retry button resumed a failed run into its own failure.
All three come from treating resume as "continue where it stopped". This post is for anyone adding human-in-the-loop or retry to a checkpointed LangGraph agent. It makes three points:
- Resuming after
interrupt()runs the whole node again from the top. Anything with a side effect goes in a node before the pause, or is made safe to repeat. - State written before a pause is still there after it. A flag that grants something, such as an approval, has to be used up explicitly, or it keeps granting.
- Resume and retry are different operations. Resume a run whose process died; start a fresh thread for a run whose agent finished with a failure.
The scripts below run offline against langgraph 1.2.11 and langgraph-checkpoint 4.1.1, the versions in my lock file at the time, plus langgraph-checkpoint-sqlite for the one test that needs a checkpoint to outlive a crashed process.
The node runs again
LangGraph documents this in the docstring of interrupt: "The graph resumes from the start of the node, re-executing all logic." It is easy to read past. Here is what it means for a node that starts paid work and then waits for approval:
submitted = [] # stands in for a paid external API
def submit_and_wait(state):
job_id = f"gpu-job-{len(submitted) + 1}"
submitted.append(job_id) # side effect BEFORE the pause
answer = interrupt({"ask": f"approve {job_id}?"})
return {"job_id": job_id, "approved": answer == "yes"}
first = app.invoke({}, cfg)
final = app.invoke(Command(resume="yes"), cfg)
paused with: {'ask': 'approve gpu-job-1?'} | submitted so far: ['gpu-job-1']
resumed, final state: {'job_id': 'gpu-job-2', 'approved': True} | submitted so far: ['gpu-job-1', 'gpu-job-2']
On resume the node starts over. interrupt() now returns the human's answer instead of pausing, but everything above it has already run a second time. The job was submitted twice, and the approval was recorded against a job the human never saw. With several interrupt() calls in one node, LangGraph matches resume values to them by their order, and the node restarts from its first line each time. A node with two interrupts, run and resumed twice:
line counts: {'top of node': 3, 'between the interrupts': 2}
The top of the node ran three times: once originally and once per resume.
The fix is structural. Put the side effect and the pause in separate nodes, so that the node that re-runs has nothing before its interrupt():
def submit(state): # side effect in its own node
job_id = f"gpu-job-{len(submitted) + 1}"
submitted.append(job_id)
return {"job_id": job_id}
def wait_for_approval(state): # the pause, with nothing before it
answer = interrupt({"ask": f"approve {state['job_id']}?"})
return {"approved": answer == "yes"}
final state: {'job_id': 'gpu-job-1', 'approved': True} | submitted: ['gpu-job-1']
Where the side effect must stay in the same node, make it idempotent, which needs a key that is the same on resume and different every other time the node runs. The node's checkpoint_ns has exactly that property: it is the node name plus a task id, and the task id was identical across both resumes above, while the same node run in three steps of a loop got three different ids. Pass it as the idempotency key to the external API, or to your own table that refuses a second submission with the same key:
key = get_config()["configurable"]["checkpoint_ns"] # e.g. 'two_questions:80fc6c6a-…'
LangGraph's functional API handles this differently: the result of a @task is stored in the checkpoint, and on resume the task is not called again. The same submit-then-interrupt flow written with @entrypoint and a @task submitted once. I haven't used the functional API beyond this probe.
The checkpoint may not have been saved
Separating the nodes assumes that once submit finishes, its result is in the checkpoint. By default that is only eventually true. durability defaults to "async", which the source describes as "persisted asynchronously while the next step executes". If the process dies in that window, the paid call happened and the checkpoint that records it does not exist.
To see it, the next script uses a SQLite checkpointer whose writes take 0.5 seconds, a submit node that records the paid call in a ledger file, and a next node that kills the process after 0.1 seconds, the way an out-of-memory error or a timeout would. Three runs per mode, then one without the delay (output condensed to one line per case):
class SlowSaver(AsyncSqliteSaver):
async def aput(self, *a, **kw): # a checkpoint store that takes 0.5 s to write
await asyncio.sleep(0.5)
return await super().aput(*a, **kw)
async def next_step(state):
await asyncio.sleep(0.1)
os._exit(1) # the process dies (OOM, timeout, deploy)
await app.ainvoke({}, config, durability=MODE)
durability=async ledger: submitted gpu-job-1 checkpoints saved: 0 (3 of 3 runs)
durability=sync ledger: submitted gpu-job-1 checkpoints saved: 3 (3 of 3 runs)
no delay, async checkpoints saved: 3
no delay, sync checkpoints saved: 3
With the 0.5-second store, async lost everything, including the checkpoint of the run's input; sync kept the result of submit every time. With the fast local store, both kept it. So this is a race, and the window is the time your checkpointer takes to write. The half second here is artificial and chosen to make the race visible. I haven't measured the loss rate on a real network store; what I'd do is measure your checkpointer's write latency, because that is the size of the window.
One detail changes how you apply this: durability is an argument to invoke and stream, not to add_node. You can't make only the paid node synchronous. The options are to run the whole graph with durability="sync" and pay one checkpoint write per step, or to keep async and record the external call somewhere durable, with an idempotency key, before making it.
An approval has to be used up
My agent asks a human before running high-risk tools. The approval was stored in a plain state channel, approval_status. The gate wrote "required" and the human's answer wrote "approved" or "rejected". Nothing ever reset it. The permission check let a high-risk call through whenever the status was "approved", so after the first approval every later high-risk call in the run passed without a question. A small graph with the same shape and two risky calls:
reset on use=False human asked about: ['send_email(to=all-staff)']
ran: ['send_email(to=all-staff)', 'delete_files(folder=/finance)']
reset on use=True human asked about: ['send_email(to=all-staff)', 'delete_files(folder=/finance)']
ran: ['send_email(to=all-staff)', 'delete_files(folder=/finance)']
The flag had been correct when approval meant resubmitting the whole job. Once approval became an interrupt() inside the run, the run continued after the answer, and nobody added the line that uses it up. It was caught before the part that lets an operator answer an approval had reached production. The fix writes "none" back on the turn that actually ran the approved call. Not on any turn the node runs: a turn where every call was skipped as a duplicate would otherwise spend the approval on nothing and send the operator back to approve a call that never happened.
A related point about the same interrupt. Whatever you put in the interrupt's value is stored in the checkpoint and returned to whoever reads the pending interrupt. My agent put the tool's arguments there so the operator could see what they were approving, and those arguments can contain tokens or signed URLs. They are now redacted before they are stored (sensitive keys masked, signatures stripped from signed URLs), while ordinary fields such as recipients and file names stay visible, because the operator still has to see what they are approving.
Resume a crash, retry a failure
In June, the Retry button in my product resumed the failed run on its existing thread. The agent woke up on the old checkpoint, read a history full of tool failures that had since been fixed, reasoned for four turns without calling a single tool, and reported failure again. The user saw a new failure; it was the old one, read back out of the checkpoint.
Resuming makes sense when the run was interrupted from outside: the process was killed, the machine went away, a deploy replaced the worker. The history in the checkpoint is still a history of work that was going well. When the agent itself concluded that the task had failed, that conclusion is in the history too, and in my case the model read it and agreed with it. Retry is now a fresh run: the same thread id with its checkpoints deleted, the user's inputs kept, and a new attempt id. A new thread id does the same job if nothing else in your system is keyed on the thread. BaseCheckpointSaver defines delete_thread and adelete_thread and the in-memory saver implements them; check that the saver you use does. Mine is a custom one and needed its own. The ownership and state checks run before the delete, because the delete can't be undone. The same reasoning is why the previous post recommends a new thread over re-invoking with a "fresh" state: the thread is the history.
Where interrupts are allowed
My resume path assumes one thing about interrupts: they only come from nodes of the top-level graph, because it keeps one list of pending questions for the whole run and routes each answer back from there. Questions raised inside a subgraph were never part of that design. A node inside a subgraph can call interrupt() without any error, and it surfaces at the top level looking like any other interrupt. That was one of the probes in the previous post. So the assumption can break quietly.
What enforced it was a lint check that parses the code and allows interrupt() to be called from three files only. That is a rule about which files, not about where in the graph the call happens, and the three allowed files contain exactly the node factories a subgraph would reuse. The check would stay green while the rule was broken. A runtime check can see the position, because each node's config carries its checkpoint_ns, and nested namespaces are joined with |. In the probe, asks is a subgraph node that calls the guard (ids shortened in the output):
def top_level_interrupt(value):
"""interrupt(), but only from a node of the top-level graph."""
ns = get_config()["configurable"].get("checkpoint_ns", "")
if "|" in ns: # nested namespaces are joined with "|"
raise RuntimeError(f"interrupt() called inside a subgraph (checkpoint_ns={ns!r})")
return interrupt(value)
outer checkpoint_ns = 'outer:c6437baa-…'
inner checkpoint_ns = 'sub:4e002875-…|inner:b8a916a6-…'
guard: interrupt() called inside a subgraph (checkpoint_ns='sub:4e002875-…|asks:64d4ff4b-…')
The separator is an internal constant of LangGraph (NS_SEP), so pin this with a test like the one above and rerun it when you upgrade.
A checklist for resumable agents
- No side effect above an
interrupt()in the same node. Split the node, put the side effect in a@task, or key it on the node'scheckpoint_ns. - Decide the durability for the whole run:
"sync", or"async"with external calls recorded durably before they are made. Measure how long your checkpointer takes to write; that is the size of the window. - Every flag that grants something is used up when it is used, on the turn that actually used it.
- Anything you put in an interrupt's value is stored and returned. Redact it before it goes in.
- Retry after an agent-reported failure starts from an empty thread (checkpoints deleted, or a new thread id) and keeps the user's inputs. Resume only after the process died.
- If your resume logic assumes interrupts come from the top level, check the namespace at run time, not the file name at lint time.
Checkpointing makes an agent resumable. It does not make every node safe to run twice, and that part is up to you.

