On every turn, my workflow agent sends Bedrock the descriptions of 99 tools: 94 from MCP servers and five for control flow, a little over 100,000 characters in all. Most workflows move through phases in a fixed order, and the phase the agent is in at any moment accepts a median of two of those tools. The other 97 are there to be picked by mistake.

The usual fix is tool search. You bind a search tool and a loader, and the model fetches the few tools it needs. I built that and measured it on four real workflows. The tool-schema part of the prompt fell by about two thirds. The runs didn't get cheaper, and the first measurement came out at 2.6 times the cost.

This is for anyone running a Claude agent with prompt caching over a large tool catalogue. After the first turn, the tool list is the cheapest part of the prompt, and a discovery step is one of the more expensive things you can add. If you want a smaller tool list, there is a way to get it without asking the model.

The catalogue in numbers

Measured today:

Tools bound on every model call99 (94 MCP tools + 5 control-flow tools)
Description text sent with them103,552 characters, about 1,101 per tool
Tools that no workflow names in any phase54 of 94
Tools a strictly ordered phase acceptsmedian 2

Every agent workflow binds the whole catalogue, because every one of them sets enable_dynamic_tool_selection, which substitutes the full manifest for the workflow's own tool list. By default each tool goes out with a generic input model, so per tool the model mostly sees a name and a description. That still comes to roughly 25,900 tokens per turn.

What tool search changes

With agent_tool_search_enabled on, a run starts with a small resident set instead of the catalogue: search_tools, load_tools, and the control-flow tools. To use a domain tool, the agent calls search_tools with a description of what it wants, calls load_tools to bind the matches, and then calls the tool. Only loaded tools reach the prompt. Each of those steps is a model call.

Run 1: a third of the schema, 2.6 times the cost

Four job-search workflows, tool search off against tool search on, same inputs. Totals across the four:

Tool search offTool search onChange
Tool-schema tokens779,511276,245−64.6%
Total context tokens1,036,9491,447,049+39.5%
Model calls32105+228%
Conversation tokens251,5111,029,510+309%
Cost (index)100257+157.3%
Workflows that finished4 of 42 of 4

"Conversation tokens" here means the part of each request that is neither tool schemas nor tool results: the messages and the tool calls themselves. That is the part discovery grows. One workflow kept searching and loading until it had made 60 model calls, and produced nothing.

Most of that was my own gate fighting the search tool

All four workflows run under a strict phase state machine. A middleware (PhaseLegalityMiddleware) rejects any tool call that isn't in the current phase's valid_next_tools, and those lists name concrete MCP tools. search_tools wasn't in any of them. The two failed runs logged this line, twice in one and four times in the other:

PhaseLegalityMiddleware rejected tool=search_tools phase=extract

The rejection message listed the tools that were legal, so sometimes the model gave up on searching and called one directly. Two workflows recovered that way. The other two kept trying to search until the run's loop and iteration guards stopped them.

The fix was an exemption list for tools that manage the agent's own work rather than acting on the user's files:

PHASE_EXEMPT_META_TOOLS: frozenset[str] = frozenset(
    {"search_tools", "load_tools", "batch_call", "read_scratchpad", "write_todos"}
)

So Run 1 measured a collision between two features of mine, not tool search. Reading the new tools against every existing gate before the first paid run would have taken ten minutes.

After the fix, it depends on what you compare

Two later measurements gave three different answers:

ComparisonTool-schema tokensCost
One workflow, alone, right after the fix−54.8%−16.4%
Full runtime, four workflows, the final run of each−72.8%+37.0%
Full runtime, four workflows, the earliest passing run of each−72.8%+1.9%

The full-runtime numbers include more than tool search. By then the runtime also had grounding checks that reject a draft whose claims aren't in the source and make the agent write it again. Those add model calls on purpose, to raise quality. The one workflow that never triggered them came out cheaper in every comparison (−8.2% on its final run). There was also a second collision: search results weren't aware of phases, so the model found tools it wasn't allowed to call yet and got rejected. Binding only the current phase's legal tools, behind another flag, fixed that. That fix swaps the tool list at every phase change, which has a cost of its own that I come back to below.

Grouped bar chart of three comparisons. Run 1 with the search tool blocked by the phase gate: schema tokens minus 64.6 percent, cost plus 157.3 percent. One workflow after the fix: schema minus 54.8 percent, cost minus 16.4 percent. Full runtime across four workflows: schema minus 72.8 percent, cost between plus 1.9 and plus 37.0 percent depending on which run is picked.
Schema tokens fell in every comparison. Cost fell in one of them.

My reading, based on the one workflow I could isolate from the quality loops: tool search on its own lands somewhere between a small saving and no change, and it never paid for itself in money. I didn't turn it on.

Why a cached catalogue makes this a bad trade

The tool list sits at the front of every request, ahead of the conversation. With a cache point after it, the first turn writes it to the prompt cache and every later turn reads it back at 10% of the input price. Removing most of it saves most of a segment that was already billed at a tenth of the input price.

A discovery step costs a whole model call: output tokens, plus a conversation that is one exchange longer for every remaining turn of the run. At the time of these runs only the prefix up to the tool list was cached, so that growing conversation was billed at the full input rate on every turn. The full-runtime measurement shows the shape clearly. Total context fell 27.3% and cache reads fell 50.8%, while conversation tokens rose 103.6%. The saving was real. It was in the cheap column, and the new cost was in the expensive one.

The context reduction does matter for one thing, which is headroom. A workflow that reads a long document has 27% more room before it hits the model's context limit. That is a reason to want a smaller tool list. It isn't a reason to make the model find the tools itself.

What about Anthropic's built-in tool search?

Anthropic now offers tool search as a server-side tool. You send every tool definition, mark most of them defer_loading, and when Claude searches, the API runs the search and expands the matching definitions inside the same response. Deferred tools stay out of the cached prefix, so caching isn't disturbed, and there are no extra round trips of the kind my two-tool version paid for.

That changes the arithmetic above, but I couldn't use it. On Bedrock it works only through the InvokeModel API, not Converse, and my agent talks to Bedrock through LangChain's ChatBedrockConverse, which uses Converse. If you're on the same stack, check that before planning around it. I haven't measured the built-in version.

Narrowing without asking the model

Every workflow already declares which tools each phase may call, because the phase middleware needs that list. So the smaller tool list can be computed once, when the workflow spec loads, as the union of its phases' valid_next_tools. No model call is involved:

declared = declared_tool_names(spec_dict)   # union of every phase's valid_next_tools
if not declared:
    return available_tools                  # prose-only phases: keep the full catalogue, log it

kept = [spec for spec in available_tools if to_langchain_name(spec) in declared]
if not kept:
    return available_tools                  # declaration matches nothing: log a warning
return kept

Of the 22 workflows that bind the whole catalogue, 18 narrow to between 7 and 12 tools (median 9, control-flow tools included). The other 4 describe their phases in prose, declare no tool lists, and keep the full catalogue with a log line saying so. For the median workflow, tool descriptions drop from about 25,900 tokens per turn to about 1,100.

Three rules keep it from doing damage, and each has tests:

  • It only removes tools. The result is always a subset of the catalogue, and a name in a workflow that no server provides drops out instead of inventing a binding.
  • It keeps the catalogue's order. The bound list is the cached prefix, and reordering it would miss the cache on every run.
  • It compares names in exactly the form the tool factory produces. If those two forms ever drift, every workflow falls back to the full catalogue while the code reports that narrowing worked. That failure is completely silent, so there is a test that checks the two forms agree for all 94 tools.

It runs once per run, never per turn. Swapping the list at each phase change, which is what fixed the second collision above, changes the bytes of the tool block and misses the cache at every phase boundary. That can cost more than the narrowing saves, which is why a later flag keeps one list for the whole run and restricts each phase through tool_choice instead. That one is also switched off.

Flow diagram. A catalogue of 94 tools plus 5 control-flow tools passes through resolve_working_set, run once at spec load with no model call, leaving 9 tools bound for the workflow at the median. A bracket under the 9 tools reads same bytes every turn, tools cache point stays valid. Three callouts: only narrows, keeps catalogue order, names must match the tool factory exactly or it falls back to all 99.
The working set is resolved once from what the workflow already declares, and the order is kept so the cached prefix doesn't change.

Where this stands: the resolver is built and switched off (agent_working_set_enabled defaults to false and no environment sets it). I haven't measured it on live runs yet. When I do, I'll judge it by cost per run and model calls, and I expect the main gain to be fewer wrong tool picks rather than a smaller bill. Anthropic's own documentation puts the point where tool selection starts to degrade at 30 to 50 available tools, and 99 is well past it.

If your workflows don't declare tools per step, the same property is available from a per-workflow allowlist written by hand. What matters is that the list is fixed before the first turn and doesn't change during the run.

What I check now before trusting a token saving

  • Cost per run and number of model calls, next to the token footprint. Never the footprint alone.
  • Completion rate. In Run 1 one workflow's cost fell 77%, because it failed after a few turns.
  • Every existing gate, read against the new tools before the first paid run.
  • More than one run per workflow, and a note of which run I compared. The full-runtime cost moved between +1.9% and +37.0% on that choice alone.