In June, a turn in my workflow builder started failing with HTTP 500 whenever the model asked for two tools at once. Bedrock's error was ValidationException: Expected toolResult blocks at messages.N.content for the following Ids. In August, my structured-extraction feature failed on a company annual report because the report did not contain one of the requested figures, and the failed run was still charged. Every local test was green in both cases. In both, the problem was in the request that LangChain builds for Bedrock, which none of my tests looked at.

This post is for anyone using ChatBedrockConverse with tools. It makes three points:

  1. The conversion from LangChain messages and tool schemas to a Converse request is part of your contract with the model. LangChain can build a request that Bedrock rejects, or one that means something different from your schema, and a mocked model never shows you either.
  2. Pin behaviour, not only versions. Run the real formatter on your real schemas in a test, and check that the deployed image installs the versions your tests ran.
  3. Decide which layer enforces what. A forced tool choice gets you a tool call, a schema shapes it, your validator keeps the constraints the schema grammar can't express, and a bad argument goes back to the model instead of out to the user.

The reproductions below run offline against langchain-aws 1.4.6 with langchain-core 1.6.0, the versions my lock file held on the post date, and against langchain-aws 1.7.4 where a newer version matters. They call private functions of langchain-aws (_format_tools, _messages_to_bedrock) to see the request without sending it. That is fragile on purpose: these are the functions whose behaviour you want a test to notice changing.

Every tool call needs an answer

The builder processes one tool per step, because each of its tools decides how the turn ends: ask the user a question, show a draft, save the result, or give up. So its tool node took the model's first tool call, ran it, and ignored the rest. That worked until the model issued two calls in one message. The next request then carried an assistant message with two toolUse blocks and a user message with one toolResult. LangChain converts this without complaint:

ai = AIMessage("", tool_calls=[
    {"name": "stage_draft", "args": {"title": "Invoice extractor"}, "id": "tooluse_A"},
    {"name": "lookup_catalog", "args": {"q": "invoice"}, "id": "tooluse_B"},
])
history = [HumanMessage("Build me an invoice workflow."), ai,
           ToolMessage("draft staged", tool_call_id="tooluse_A")]   # B never answered

messages, _system = _messages_to_bedrock(history)
toolUse ids   : ['tooluse_A', 'tooluse_B']
toolResult ids: ['tooluse_A']
unanswered    : ['tooluse_B']

Bedrock does complain. Every toolUse id in an assistant message needs a matching toolResult in the next user message, and a request without one is rejected outright. My fix kept the one-tool-per-step rule and answered every other call with a result that says so:

ToolMessage(
    content=json.dumps({"deferred": "Only one builder tool is processed per step; this call "
                                    "was not run. Re-issue it on a later turn if still needed."}),
    tool_call_id=tc["id"],
    name=tc["name"],
)

The wording matters as much as the id. A result that says what happened lets the model decide whether to ask again; an empty string or a fake success would teach it something false about the tool.

LangGraph's prebuilt ToolNode runs every call in the message, so this bites custom tool nodes: anything that indexes tool_calls[0], drops a call it considers a duplicate, or skips a call a policy rejected. Each of those still needs to put a ToolMessage with that call's id into the history.

The formatter changed my schema

The extraction feature takes a JSON Schema from the caller, binds it as a tool, and validates the model's answer against the same schema afterwards. The failing case came from an evaluation set run against the deployed service: an annual report and a schema with "auditor_fee_total": {"type": ["number", "null"]}. Annual reports of that kind don't disclose audit fees, so the correct answer is null. The model answered '<UNKNOWN>', was asked to repair, answered the string 'null', and the run failed validation both times.

Why would the deployed service behave differently from the tests? The Lambda image installed dependencies from a requirements file with version ranges; every local gate installed from poetry.lock. Resolving the requirements file for the image's platform, without building the image, shows what it would actually install:

uv pip compile requirements-lambda.txt --python-version 3.12     --python-platform aarch64-manylinux2014

Compared with the lock:

Packagepoetry.lock (tests)Image (production)
langchain1.3.161.3.18
langchain-core1.6.01.6.1
langchain-aws1.4.61.7.0
langchain-openai1.2.21.6.0
langchain-anthropic1.5.61.7.0

The langchain-aws gap was the one that mattered. The way to find it was to compare the three places the schema appeared in that one run: the prompt and the validator both carried ["number", "null"], while the toolSpec sent to Bedrock carried "number". When a model seems to ignore your schema, check which copy of the schema it was actually given. Version 1.7 removes the null branch from every union in a tool schema before sending it. Here are three nullable shapes through both versions' _format_tools:

langchain-aws 1.4.6
  Optional[str]  {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'default': None}
  fee (required) ['number', 'null']
  table cell     ['string', 'number', 'null']
langchain-aws 1.7.4
  Optional[str]  {'type': 'string', 'default': None}
  fee (required) number
  table cell     ['string', 'number']

The change is deliberate. The function's docstring in 1.7.4 says Converse allows at most 16 union-typed parameters across all tools in a request, and that optional MCP parameters use up that budget. For an optional property that is a fair trade: the model can leave the property out, and the schema never needed a null. It breaks two cases where leaving out is not possible. A required field that may be null can no longer be null. And an array item can't be skipped without shifting every later element, so a blank table cell has nothing it is allowed to be.

This matches what the evaluation set showed. Three other cases on the same deployed image declared a nullable tax_id that was not required, and they passed because the model left the key out; the narrowing was there too, just harmless. In the failing case the field was required, so leaving it out was not an option. The schema the model received didn't allow null, the validator would have accepted null, and an honest "not disclosed" became a failed, billed run. The union is not a Bedrock limitation, at least not with a handful of fields: one call to Claude Haiku 4.5 on Bedrock with {"type": ["string", "null"]} in the tool schema was accepted and answered {"vendor": "Northwind Retail", "tax_id": null}.

Two pins: the version, and the behaviour

The first fix pinned the five packages to their lock versions in the image's requirements, with a test that compares every entry the two files share. That guarantees the image runs what the tests ran. It says nothing about whether that version is right: the obvious next step, poetry update langchain-aws, moves the lock to 1.7, and from then on the same test insists that the image install 1.7 as well. The defect comes back with both checks green.

So the second fix, the one I would copy first, pins the behaviour. It runs the installed formatter over the shapes you rely on and asserts what Bedrock would receive:

def sent_properties(tool) -> dict:
    """What Bedrock would receive for this tool, via the installed formatter."""
    return bedrock_converse._format_tools([tool])[0]["toolSpec"]["inputSchema"]["json"]["properties"]

def test_a_json_schema_null_union_survives():
    tool = {"type": "function", "function": {"name": "emit", "description": "Return the fields.",
            "parameters": {"type": "object", "required": ["fee"],
                           "properties": {"fee": {"type": ["number", "null"]}}}}}
    assert sent_properties(tool)["fee"]["type"] == ["number", "null"]

def test_a_pydantic_optional_survives():
    class Invoice(BaseModel):
        """Return the extracted fields."""
        tax_id: Optional[str] = None
    assert {"type": "null"} in sent_properties(convert_to_openai_tool(Invoice))["tax_id"].get("anyOf", [])
langchain-aws 1.4.6:  2 passed
langchain-aws 1.7.4:  2 failed

The test catches the change whichever version the lock names. My own version adds two guards against a vacuous pass: one checks that the formatter really wraps its input in a toolSpec, and one checks that a plain non-nullable field stays a bare type.

If you need 1.7 or later

Pinning 1.4.6 is a holding position, not an answer. Three ways forward, depending on the field:

  • If the field can be optional, make it optional and treat a missing key as null in your own validation. That is what the narrowing assumes, and it is why the three optional-field cases above kept passing.
  • If the field must be required and nullable, or it is an array item that may be blank, pass the native Converse toolConfig yourself with llm.bind(toolConfig={"tools": [{"toolSpec": ...}]}). Handing the same toolSpec dict to bind_tools does not help: it is converted back to the generic format and stripped when the request is built. Comparing the request each path produces on 1.7.4:
    bind_tools([toolSpec])   {'type': 'number'}
    bind(toolConfig=...)     {'type': ['number', 'null']}
    With your own toolConfig you also own toolChoice, any cache point, and the 16-union budget the stripping was protecting, so count your unions if you bind many tools.
  • Keep the behaviour test either way. It is what tells you which of these you still need after the next upgrade.

Forcing a tool is not enforcing a schema

To make the extraction answer in the caller's shape, the schema is bound as the only tool with tool_choice="any". That guarantees the model answers with a call to that tool. It does not by itself guarantee the arguments match the schema. What made the difference in practice was also closing every object with additionalProperties: false before binding: before that change, a request for a single property came back with eleven.

A full guarantee needs strict tool use ("strict": true), where the output is constrained by a grammar built from the schema. Anthropic's structured outputs documentation lists what that grammar can't express, including minimum, maximum, minLength, maxLength, pattern and recursive schemas, and requires additionalProperties: false. My code strips those keywords from the schema before binding it. In langchain-aws, bind_tools(..., strict=True) sets the flag on each toolSpec. My own strict path was behind a flag that was off at the time, so I haven't measured it. If your model supports structured outputs on Bedrock (the same page lists which Claude models on Bedrock support it), strict is the option designed to guarantee the shape, though I can't give you a measurement; the validator stays either way.

Either way, the constraints the grammar drops are still the caller's contract, so the validator checks the full, unprojected schema. The split is: the binding decides the shape, the validator decides the constraints. One test sends {"n": 5} against {"minimum": 10} and expects a rejection; a second asserts that the bound schema has no minimum in it, because otherwise the first test could pass for the wrong reason.

A bad argument should reach the model, not the user

So far the contract has run from your code to Bedrock. It also runs the other way, from the model's output into your tools. In July, a builder turn ended in HTTP 500 because the model wrote a 1,500-character reason into a tool whose argument schema said max_length=1024. StructuredTool.ainvoke validates arguments before running the tool and, by default, raises. The node calling it had no handler, so the error left the graph. LangChain has a setting for this, handle_validation_error:

default : ValidationError raised out of ainvoke
handled : ToolMessage 'Tool input validation error'
callable: ToolMessage 'Invalid arguments: reason: String should have at most 1024 characters'

True turns the error into a tool message, but a generic one the model can't act on. A function gives the model the field and the rule it broke. My own fix did two things instead: the tool clamps an over-long reason itself, because truncating it loses nothing that matters, and the calling node catches any remaining validation error and ends the turn with a normal "could not build this" answer instead of an error. The general rule I follow: repair it in code when the repair is obvious, return a specific error to the model when the model can fix it, and never let it reach the user as a 500.

Test doubles that skip the contract

All of these bugs live between your code and the provider, which is exactly the layer a mocked chat model replaces. One example from the same August change: the test fixture's model was a MagicMock, and the code under test called model.bind_tools(...) and then ainvoke on the result.

bound is model: False
model.ainvoke awaited: 0 | bound.ainvoke is a MagicMock
after fix, bound is model: True

MagicMock.bind_tools returns a new child mock, so every assertion about model.ainvoke checked an object the code never called. Without model.bind_tools.return_value = model, 23 tests in that file failed: they made assertions about model.ainvoke, which the code never called, because it called the child mock instead. The mock is not wrong to use. What it can't do is tell you anything about the request, so keep at least one test per rule in this post that runs the real langchain-aws formatting offline.

A checklist for your own agent

  1. Every tool call in an assistant message gets a ToolMessage with its id, including calls you skip, defer or reject.
  2. The deployed image installs the same versions your tests ran. Resolve the image's requirements for its platform and compare against the lock in a test.
  3. For each schema feature you depend on (nullable fields, enums, nested arrays), a test runs the installed formatter and asserts the feature reaches the toolSpec.
  4. Forced tool choice for the shape, additionalProperties: false on every object, and a validator on the full schema for the constraints the provider's grammar drops.
  5. Tool argument errors are handled inside the graph: repaired in code, or returned to the model with the field and the rule.
  6. At least one test per rule above runs without mocking the chat model's request formatting.

This is Bedrock-specific in its details. The toolResult rule, the union budget and the strict plumbing are Converse behaviour. The shape of the problem is not: whatever provider you use, the adapter between LangChain and its API is code you ship, and it deserves tests that run it.