My workflow-building agent has six internal tools: list the available tools, propose phases, check a phase against the tool catalog, draft a rubric, estimate cost, compose the prompt. (A seventh, for retrieval, has its own adapter and isn't part of this.) They've always run as in-process stand-ins that return simplified answers. Since June there's been a wrapper that sends them to a real MCP server instead, behind a switch that stayed off, and it fails open: if the server call doesn't succeed, the tool quietly returns the stand-in's answer.

This week I switched it on. Fail-open hid three separate problems along the way, each of which would have looked like a working agent, and I kept it anyway. If your agent's tools cross a boundary and fall back when it fails, the fallback makes "every call rejected" look the same as "every call fine", so something other than the runtime has to tell them apart. For me that's a parity test in CI and, after a switch, reading the server's logs until the fallback count is zero.

Why fail open at all

These tools are read-only and advisory. They suggest, check and estimate; the agent decides. The fallback is what lets a conversation finish when the tool server really is down, and failing closed would turn an outage into a stuck conversation. So the wrapper logs a warning and returns the stand-in's result on any unsuccessful call, rejections included, and neither the user nor the model can tell.

I wrote down at the time that this means the guarantee can only live in CI: the runtime will never catch a mismatch between the two sides, by design. I also chose not to turn the warning into an error log: the wire path wasn't even reachable yet, and nobody would have been reading that log for a defect a test already pinned.

What it hid

A shape mismatch. The server's input models reject any field they don't declare. The phase checker's stand-in sent phases[{name, tool_id}]; the server expected phases[{phase_id, tools}]. Every call would have been rejected and answered by the stand-in, which, being a stand-in, said the phases were fine. I had a parity test already, from six weeks earlier, and it had found five of the six tools sending top-level fields the server rejected. It compared top-level names, and by that measure the phase checker had always been the one tool that matched: both sides took a field called phases. The mismatch was one level down.

An id mismatch. The agent names tools <server>-mcp:<tool>. The server's index of known tools used a second form, <server>__<tool>, the name the runtime routes on internally. Every real tool id would have come back "not found", which is an unsuccessful call, so the stand-in would have answered; the cost estimator on the same index would have priced every tool at zero. The phase proposer had the mirror image: it returned internal names, so its suggestions never matched the agent's list of allowed tools. No schema comparison sees this, since both sides are just strings.

A missing directory. With the switch on in development, the logs of the first real turn showed the server answering five of six calls with DEPENDENCY_NOT_FOUND, its error when a data file lookup fails, and all five quietly handed to the stand-ins. I first read that as the agent guessing bad ids, but even a phase template that does exist (the proposer builds phases from templates on disk) was rejected. The server's container image was built from its source directory only; the data files these tools read had never been copied in, in any environment. The conversation looked fine. I found it in the server's logs, not the chat.

What catches it instead

The new parity test reduces both sides to the same thing: a set of field paths like phases[].phase_id. For the agent side it uses the stand-in's argument schema, plus the keys the stand-in actually returns when you call it once. For the server side it uses the input and output schemas the server dumps from its own models into a manifest at build time. It records two kinds of difference per tool: in: paths the agent sends that the server would reject, and out: keys the agent was promised that never arrive. The second kind never raises anything; the model reads a field the description mentioned, finds nothing, and carries on. A minimal repro of the input side: the old check, the new check, and a fail-open wrapper.

top-level names the server lacks: set()
field paths the server lacks:     ['phases[].name', 'phases[].tool_id']
{'source': 'stub', 'ok': True}
{'source': 'stub', 'ok': True}
{'source': 'stub', 'ok': True}
fallbacks: 3 of 3
The probe (Python 3.11, pydantic 2.13.5)
from pydantic import BaseModel, ConfigDict, ValidationError
from pydantic_core import to_jsonable_python

class Strict(BaseModel):
    model_config = ConfigDict(extra="forbid")
class Binding(Strict):            # server side
    phase_id: str
    tools: list[str]
class ServerInput(Strict):
    phases: list[Binding]
class Phase(BaseModel):           # agent side (the old in-process stub)
    name: str
    tool_id: str
class StubArgs(BaseModel):
    phases: list[Phase]

def paths(model):                 # no anyOf / Optional[...] handling: enough for this probe
    schema = model.model_json_schema(); defs = schema.get("$defs", {})
    def walk(node, prefix):
        node = defs[node["$ref"].split("/")[-1]] if "$ref" in node else node
        if node.get("type") == "array":
            return walk(node["items"], prefix + "[]")
        props = node.get("properties")
        if not props:
            return {prefix}
        return {p for k, v in props.items() for p in walk(v, f"{prefix}.{k}" if prefix else k)}
    return walk(schema, "")

print("top-level names the server lacks:", set(StubArgs.model_fields) - set(ServerInput.model_fields))
print("field paths the server lacks:    ", sorted(paths(StubArgs) - paths(ServerInput)))

fallbacks = 0
def call_tool(args: StubArgs):
    global fallbacks
    try:
        return {"source": "server", **ServerInput.model_validate(to_jsonable_python(args)).model_dump()}
    except ValidationError:
        fallbacks += 1
        return {"source": "stub", "ok": True}            # fail-open: canned answer
for _ in range(3):
    print(call_tool(StubArgs(phases=[Phase(name="parse", tool_id="pdf:extract")])))
print("fallbacks:", fallbacks, "of 3")

The differences are kept in a baseline that can only shrink. Measured the day before the switch, all six tools were in it: the phase checker for inputs and outputs, the other five for output keys. A new difference fails the test, so does an entry left behind after its tool is fixed, and the switch was turned on in development only once the baseline was empty. The id mismatch got its own tests that send a real id from one side to the other, and the server's index is now keyed on the agent's form. The container image now includes the data directory, with a test that it lands where the code looks.

Then the logs. I ran a real conversation one turn at a time after the switch and counted fallbacks in the server's logs. The fourth turn made four calls to the server, fell back zero times, and ended with a draft. The warning is still just a warning; counting it is a step in turning the switch on, done by hand, not something the system watches for me.

What fail-open hidWhy the runtime couldn't see itWhat caught it
Nested field mismatchRejected call, stand-in answeredField-path parity test in CI
Tool id format mismatch"Not found", stand-in answeredTests that send a real id across
Server image missing its data filesEvery data-backed call failed, stand-ins answeredReading the server's logs after the switch; now a test on the image

One failure the fallback didn't hide: the first turn also returned HTTP 500, because LangChain's StructuredTool hands the wrapper pydantic objects, the client that sends the call serializes with json.dumps, and the resulting TypeError wasn't the kind of error the fallback catches. Converting the arguments with to_jsonable_python first fixed it.

Where I'd fail closed

This trade only works for tools whose canned answers are harmless, like these advisory ones, and only if the CI check is real. For a tool that writes data or spends money I'd fail closed and take the stuck conversation. And all of this comes from one migration of six tools and one test conversation after the switch.