On September 1 a subtitle-generation run on my dev stack stopped at turn 8 with token_budget_exceeded: 384064 > 350000. The agent hadn't done much by then. Its conversation had grown by 15,584 tokens over eight turns. The other 368,480 tokens were one 46,060-token prompt prefix (the system prompt plus the tool catalogue), served from the prompt cache on every turn and counted again on every turn.

The next day I found the same mistake in the cost calculation, which matters more, because that number is the run's cost. input_tokens means different things depending on which layer reports it. LangChain's version includes cached tokens. The pricing model I use, which follows Anthropic's, assumes it doesn't. My code passed LangChain's number into that model next to the separate cache counts, so each cached token was priced twice: once at the full input rate and once at the cache-read rate.

If you meter Claude usage through LangChain, or any layer that normalizes provider responses, and price it from a rate table, this is worth ten minutes of checking. The short version: find out which convention your token counts follow, then list every piece of code that reads them, because they are usually asking different questions.

Two conventions for one field name

Anthropic's Messages API reports input as three fields that don't overlap. The prompt caching docs define input_tokens as the tokens "which were not read from or used to create a cache", and give the total as:

total_input_tokens = cache_read_input_tokens
                   + cache_creation_input_tokens
                   + input_tokens

LangChain's chat-model integrations report usage as a UsageMetadata dict. In langchain_core, input_tokens is documented as "Sum of all input token types", and the cache figures sit under input_token_details as a breakdown of that sum. Here is turn 2 of the failing run, as the agent node received it:

{
    "input_tokens": 47_084,        # includes the cached prefix
    "input_token_details": {
        "cache_read": 46_060,      # part of the 47,084, not in addition to it
        "cache_creation": 0,
    },
}

Neither convention is wrong. The bug appears when a number produced under one is consumed by code written for the other. My rate table and the function that applies it, compute_llm_cost_micro_u, use Anthropic's shape: three exclusive categories, each with its own rate. On Bedrock a cache read costs 10% of the input rate and a five-minute cache write costs 125%.

What the double count did to one turn

Here is that same turn priced both ways, in input-token equivalents so the rates are easy to follow:

What the formula receivedInput-token equivalents
Priced as intendedinput 1,024 × 1.0 + cache read 46,060 × 0.15,630
Priced before the fixinput 47,084 × 1.0 + cache read 46,060 × 0.151,690

That is 9.2 times the real input-side cost. Across the seven turns of the run that read the cache, the ratio ran from 7.3× to 9.2×. Over the whole run's input side it came to 4.5×, pulled down by turn 1, which wrote the cache instead of reading it. These figures leave out output tokens, which were priced correctly and aren't in the data I'm quoting.

The per-turn cost also feeds the per-run spending cap, so a cache-heavy run would reach that cap long before its real spend justified it.

The token limit had the same bug

The September 1 run wasn't stopped by the spending cap. It was stopped by a separate limit on cumulative input tokens, computed per workflow as max(100_000, max_iterations × 10_000). The subtitle workflow declares 35 iterations, so its ceiling is 350,000.

That limit read LangChain's gross input_tokens as well. With a 46,060-token prefix re-counted every turn, each turn cost the limit about 48,000 tokens regardless of what the agent did. 350,000 divided by 48,000 is a little over 7. The workflow declared 35 turns and could run 7.

Line chart over eight turns. Cumulative gross input_tokens rises in a straight line from 46,564 to 384,064 and crosses a dashed ceiling at 350,000 on turn 8, where the run stopped. Cumulative new input stays near the axis and ends at 15,584.
Cumulative input per turn for the failing run. The gross counter crossed the ceiling at turn 8; new input across all eight turns was 15,584 tokens.

The constants weren't wrong when they were chosen. The March commit that introduced the formula sized it for about 2,800 tokens per iteration. The per-turn tool catalogue that makes up most of today's prefix didn't arrive until May. The threshold stayed where it was while the thing it measured grew about 17 times.

One count, three questions

Once I listed every place that read the token count, most of the fix was deciding which question each one was asking:

ReaderQuestionNumber it needs
Cost formula and spending capWhat did this turn cost?Non-cached input, plus cache reads and writes at their own rates
Cumulative token limitHow much has this conversation grown?Non-cached input only
Telemetry and logsHow many tokens did the provider process?The gross count

The gross number is the right answer for the third row, so it stayed there. Telemetry wants the total the model read on that turn, cached or not, because that is what latency and context-window headroom track.

The fix

One helper computes the non-cached part, and the two readers that need it call the helper:

def fresh_input_tokens(usage: Any) -> int:
    """Input tokens this turn that were NOT served from the prompt cache."""
    total = extract_token_count(usage, "input_tokens")
    cache_read, cache_write = extract_cache_tokens(usage)
    return max(0, total - cache_read - cache_write)

Two details are deliberate. A provider that reports no input_token_details gets (0, 0) for the cache pair, so for it the result is exactly input_tokens and nothing changes. And the floor at zero means a provider that ever reports a breakdown larger than its total can't hand a run negative usage, which would make the limit unreachable.

The cost path now passes fresh_input_tokens(usage) as input_tokens, next to the cache counts it already passed. That change has no switch. I didn't want a way to go back to pricing from the wrong number.

The token limit changed differently. Its threshold is one of the deterministic limits I don't let anything adjust casually, so the constants stayed exactly as they were. What changed is which counter the limit reads, behind a switch (agent_token_gate_cache_aware_enabled) that defaults off in code and is set on for both environments in the infrastructure definition. Both counters are recorded on every turn whatever the switch says, and the end-of-run summary prints both, so the re-counted prefix shows up as the difference between them in every log.

The sentence that hid the second bug

I fixed the token limit first. That change's description said money was still bounded by the spending cap, which "prices cache reads and writes properly". I believed it because the cap's code did read the cache counts. I hadn't looked at what it read next to them.

The same change did open a separate issue to measure whether the cost path double-counted, rather than editing it on a hunch. The measurement came back the next day and said it did. The habit I'm keeping from this: "X is still protected by Y" is a claim about Y's inputs, and it needs the same check as the code it's reassuring me about.

Tests built from the failing run

Both regression tests use the run's real numbers. The cost test doesn't only assert that the new price is lower. It asserts that the old overcharge was exactly the cached tokens priced a second time at the input rate:

usage = {"input_tokens": 47_084, "input_token_details": {"cache_read": 46_060}}
fresh_inp = fresh_input_tokens(usage)
assert fresh_inp == 1_024

correct_cost = _compute_turn_cost_micro(..., input_tokens=fresh_inp,
                                        cache_read_tokens=46_060, cache_write_tokens=0)
buggy_cost   = _compute_turn_cost_micro(..., input_tokens=47_084,
                                        cache_read_tokens=46_060, cache_write_tokens=0)
overcharge_only = _compute_turn_cost_micro(..., input_tokens=46_060,
                                           cache_read_tokens=0, cache_write_tokens=0)

assert buggy_cost - correct_cost == overcharge_only
assert buggy_cost > correct_cost * 5

Cost is linear in the rates, so that identity holds whatever the price table says. The test survives the next rate update and fails only if someone prices the gross count again. It also has a gap I haven't closed: it calls _compute_turn_cost_micro directly rather than driving the agent node, so a regression at the call site in reason.py would get past it. Nothing yet runs the node with a cached usage block.

The token-limit test file holds all eight turns. One test asserts that the gross total of those turns still exceeds 350,000, and only then does another assert that the fixed limit lets the run through. Without the first, the second could pass on a fixture that never tripped anything in the first place.

Checking your own stack

Log one raw usage block from a turn that read a large cached prefix and added only a short message, which is most turns of a tool-using agent. Then look at input_tokens:

  • If it is roughly the size of the new message and far smaller than cache_read, your layer uses the exclusive convention. Anthropic's raw API does.
  • If it is at least cache_read + cache_creation, and subtracting those two leaves roughly the size of the new message, it uses the inclusive one. LangChain's UsageMetadata does.

A turn with a small cache hit won't tell you much, because under the exclusive convention the new content can outweigh the cached part.

Then find every reader of that field and write down which of the three questions it answers. In my case two readers, the cost formula and the token limit, needed a different number from the one they were given. Everything else was right to keep the gross count.