In early September, my product offered a free Retry whenever a run stopped at one of four limits, such as a token budget. For two of the four, the retry failed on its first turn before doing any work. It re-ran the job on the same LangGraph thread with a fresh initial state, and the token counter that had tripped the limit was still at 384,064 when the retried node read it. A week later, while checking whether my agent could be split into several nodes that each see only part of the state, I found that a node declared to read one field can write every field in the graph.
Both come from the same property of LangGraph state, and this post is for anyone splitting an agent into several nodes or subgraphs over one shared state. It makes three points:
input_schemalimits what a node reads, not what it writes. Any node can overwrite any channel of the graph's state, the last write wins, and a key that isn't a channel is dropped without an error. If a channel has an owner, you have to enforce that yourself.- Every write goes through the channel's reducer, including the input you pass to a new run on an existing thread. A "fresh" initial state does not reset an additive counter. Use
Overwrite, or start a new thread. - A subgraph writes its checkpoints into the parent's checkpointer, under its own
checkpoint_ns, even if you compile it with a checkpointer of its own. A custom checkpointer that ignorescheckpoint_nsis not safe to use with subgraphs.
Every output below comes from scripts that run offline against langgraph 1.2.11 and langgraph-checkpoint 4.1.1, the versions in my lock file at the time. The original measurement also re-ran on checkpoint 4.2.0 with identical results.
Why I tested this instead of reading it
A design document proposed splitting my agent's single LangGraph loop into several nodes, each with a narrow view of the state, and claimed LangGraph provides that isolation natively. The claim came from reading the source, so I ran the questions it depended on in a script anyone can rerun. Two answers were not what the design assumed.
Reads are narrowed, writes are not
The probe has a state with four channels and two nodes. narrow_node is added with input_schema=NarrowIn, which declares only a, and deliberately returns two channels it did not declare. wide_node reads everything and writes one of the same channels:
class Full(TypedDict, total=False):
a: str
b: str
singleton: str
log: Annotated[list[str], operator.add]
class NarrowIn(TypedDict, total=False):
a: str
def narrow_node(state):
return {"b": "written-by-narrow", "singleton": "narrow-wins", "log": ["narrow ran"]}
def wide_node(state):
return {"singleton": "wide-wins", "log": ["wide ran"]}
g.add_node("narrow", narrow_node, input_schema=NarrowIn)
g.add_node("wide", wide_node) # START -> narrow -> wide -> END
state keys the NARROW node actually saw : ['a']
state keys the WIDE node actually saw : ['a', 'b', 'log', 'singleton']
final b : written-by-narrow
final singleton : wide-wins
The read side works as documented: the narrow node saw only a. The write side is not restricted at all. The narrow node's write to b landed, and singleton, a channel without a reducer, holds whatever the last node wrote. In this graph the order is fixed, so "last" is predictable. When two nodes write the same plain channel in one super-step, LangGraph does raise:
InvalidUpdateError At key 'singleton': Can receive only one value per step. Use an Annotated key to handle multiple values.
Across steps it doesn't. The later write replaces the earlier one without any error, which is the case you are in whenever nodes run one after another.
There is a second way to lose a write. A key that is not a channel at all is dropped:
g.add_node("n", lambda s: {"documnet": "typo", "extra": 1}) # misspelt key + unknown key
print(g.compile().invoke({"document": "q3 report"}))
{'document': 'q3 report'}
No error, no warning. A misspelt state key looks exactly like a node that decided not to update anything.
Why this mattered for my agent: several channels in its state are single values that a guard reads. One records the last tool call, and the phase logic derives the current phase from it. One records whether a human approved a risky action. Two are counters that stop a run that has looped too long. If several nodes can write all of these, any one of them can move the phase, grant an approval, or reset a counter, and nothing in the framework will notice.
Enforcing ownership yourself
add_node in 1.2.11 has input_schema but no per-node output schema, so ownership has to be your own code. A decorator that declares what a node may write and fails loudly otherwise, for sync and async nodes:
def writes(*allowed):
"""Declare which channels a node may write; fail loudly on anything else."""
def check(name, update):
keys = set(update.update or {}) if isinstance(update, Command) else set(update or {})
extra = keys - set(allowed)
if extra:
raise ValueError(f"{name} wrote undeclared channel(s): {sorted(extra)}")
return update
def wrap(fn):
if inspect.iscoroutinefunction(fn):
async def node(state):
return check(fn.__name__, await fn(state))
else:
def node(state):
return check(fn.__name__, fn(state))
node.__name__, node.__writes__ = fn.__name__, allowed
return node
return wrap
def summarize(state): # a node that "helpfully" sets a flag it doesn't own
return {"document": state["document"].upper(), "approval_status": "approved"}
unguarded: {'document': 'Q3 REPORT', 'approval_status': 'approved'}
guarded : summarize wrote undeclared channel(s): ['approval_status']
It checks Command.update too, for nodes that route with Command. The wrapper takes only state; a node that also needs config or a store needs those parameters added to the wrapper, because LangGraph decides what to inject by reading the node's signature. The same check catches the misspelt key, because documnet is not in the declared set.
Two tests make it hold over time. One fails if a node has no declared write set, so a new node can't skip the declaration. The other fails if a channel without a reducer has more than one declared writer, which turns "who owns approval_status?" into a question with one answer:
def declared_writes(spec):
fn = getattr(spec.runnable, "afunc", None) or getattr(spec.runnable, "func", None)
return getattr(fn, "__writes__", None)
missing = [name for name, spec in builder.nodes.items() if declared_writes(spec) is None]
assert not missing, f"nodes without a declared write set: {missing}"
writers = defaultdict(list)
for name, spec in builder.nodes.items():
for channel in declared_writes(spec) or ():
writers[channel].append(name)
shared = {ch: w for ch, w in writers.items()
if len(w) > 1 and not isinstance(builder.channels[ch], BinaryOperatorAggregate)}
assert not shared, f"channels with several writers and no reducer: {shared}"
On a graph where two nodes both declare approval_status, the second test reports {'approval_status': ['b', 'c']}.
A new run's input goes through the reducers
Back to the Retry button. The free retry re-invoked the failed job on its existing thread, passing the same initial state a new job starts with, on the assumption that this would start the run from a clean state. LangGraph doesn't overwrite the checkpointed state with your input. It applies the input to each channel through that channel's reducer, the same way it applies a node's return value. Reduced to the three kinds of counter involved:
def merge_dicts(a: dict, b: dict) -> dict:
return {**a, **b}
class State(TypedDict):
input_tokens: Annotated[int, operator.add] # additive counter
reads: Annotated[dict, merge_dicts] # per-key counter
iterations: int # no reducer: last write wins
def work(state): # the values the real run ended with
print(f" node sees: input_tokens={state['input_tokens']}, ...")
return {"input_tokens": 384_064, "reads": {"__total__": 20}, "iterations": 35}
fresh = {"input_tokens": 0, "reads": {}, "iterations": 0}
await app.ainvoke(fresh, cfg) # first run
await app.ainvoke(fresh, cfg) # "retry" on the same thread
await app.ainvoke({"input_tokens": Overwrite(0), "reads": Overwrite({}), "iterations": 0}, cfg)
first run
node sees: input_tokens=0, reads={}, iterations=0
retry: same thread, fresh state as input
node sees: input_tokens=384064, reads={'__total__': 20}, iterations=0
retry: same, with Overwrite
node sees: input_tokens=0, reads={}, iterations=0
Adding zero to 384,064 leaves 384,064, and merging an empty dict changes nothing. Only the plain integer reset. In my product the four limits split exactly along that line: the iteration cap and the repeated-call cap were plain integers, so the retry cleared them; the token budget and the read budget used additive reducers, so the retry hit the same limit before doing anything. Nothing anywhere in the code reset the token counter; only a full rerun, which deletes the thread's checkpoints, did.
For those two limits my product now stops offering the free retry. The general rule: if a retry means "start again", either start a new thread, or reset each reducer channel explicitly with langgraph.types.Overwrite, which bypasses the reducer for that one write. Don't rely on passing a fresh state, and don't decide which counters survive by reading the code; run the retry once in a test and look at what the first node sees.
You don't have to maintain the list of reducer channels by hand. The graph builder knows which channels have one, so a retry input can be built from it, keeping whatever the retry should carry over, such as the message history:
reducer_channels = [k for k, ch in builder.channels.items()
if isinstance(ch, BinaryOperatorAggregate)]
KEEP = {"messages"} # what a retry should carry over
retry_input = {k: (Overwrite(v) if k in reducer_channels else v)
for k, v in fresh.items() if k not in KEEP}
channels with a reducer: ['messages', 'input_tokens', 'reads']
retry input: {'input_tokens': Overwrite(value=0, ...), 'reads': Overwrite(value={}, ...), 'iterations': 0}
Subgraphs write into the parent's checkpointer
The other way to split an agent is into subgraphs. The probe compiles a small subgraph three ways and runs it inside a parent graph that has an InMemorySaver, then counts what is stored. The counts depend on how many steps each graph has; what matters is the second namespace (thread labels and namespace ids shortened):
sub compiled with checkpointer=None (default)
parent store thread='t-None (default)' ns='' -> 4 checkpoint(s)
parent store thread='t-None (default)' ns='sub:3bef…' -> 3 checkpoint(s)
sub compiled with checkpointer=its own InMemorySaver()
parent store thread='t-its own …' ns='' -> 4 checkpoint(s)
parent store thread='t-its own …' ns='sub:e94d…' -> 3 checkpoint(s)
sub compiled with checkpointer=False
parent store thread='t-False' ns='' -> 4 checkpoint(s)
The subgraph's checkpoints go into the parent's store, on the parent's thread, separated only by checkpoint_ns. Giving the subgraph its own saver changes nothing; in a separate run, that saver was still empty afterwards. The configuration reaches the subgraph through a context variable, and the parent's checkpointer comes with it. checkpointer=False is the only option that opts the subgraph out. The documented checkpointer=True, which gives a subgraph state that persists across calls, also writes into the parent's store, under a stable namespace (sub instead of sub:<task id>).
For LangGraph's own savers this is fine, because they store and query by checkpoint_ns. My production saver is a custom DynamoDB one, keyed on (thread_id, checkpoint_id). It passes checkpoint_ns through, but the namespace is not part of the key and not part of any query. Its "latest checkpoint for this thread" query could therefore return the subgraph's checkpoint when the parent asked for its own. Nothing went wrong in production, because nothing used subgraphs yet; the probe found it before anything did. If you have written your own checkpointer, check this before your first subgraph.
Two smaller results from the same script are worth knowing. Send can carry a payload whose shape differs from the main state, including keys that are not channels of the main graph, which is what you want for a map step. And a node inside a subgraph can call interrupt(), which surfaces at the top level; that one has consequences for resuming, and it is part of the next post.
A checklist before splitting an agent
- List every channel and write down which node owns it. A channel with more than one writer needs a reducer that makes the combination correct, or a single owner.
- Give each node a declared write set and fail on anything outside it. That also catches misspelt keys, which LangGraph drops without an error.
- For every retry or restart path, run it once in a test and assert what the first node sees. Reset reducer channels with
Overwrite, or use a new thread. - If you use a custom checkpointer, confirm that it stores and filters by
checkpoint_nsbefore you add a subgraph. Otherwise compile the subgraph withcheckpointer=False. - Rerun the probes when you upgrade
langgraph. None of these behaviours is pinned by a test of mine, and any of them could change.
None of this is a defect in LangGraph. Shared, reducer-merged state is the design, and it is what makes fan-out and message accumulation easy. The mistake is assuming that a narrower view of the state is also a narrower set of permissions.

