How the MCP compatibility tests work

A result on these pages means a client was observed doing something, not that a vendor said it could. Here is how that observation is made and what each status does and does not tell you.

The setup

Every test points a real MCP client at a small test server that the lab controls. The server offers known fixtures: one tool, one prompt, and one resource. Sitting between the server and its transport is a recorder that writes down every JSON-RPC message in both directions, with a timestamp and a sequence number, before the server acts on it.

The client is then asked to use one feature. For an automated client like Claude Code, the lab runs it headlessly with only the test server configured. For a client that needs a person at the keyboard, an operator follows written instructions and the lab records the same traffic either way.

Why the wire decides

A status is decided by what the server saw, not by what the client or its model says afterwards. If a test expects a tools/call naming the fixture tool, the result is supported only when that exact request shows up in the recording. The model's final answer is kept as supporting evidence, because a model can describe a tool call it never made.

Each run gets a fresh random nonce, and the client is asked to pass it along, for example as a tool argument. A request only counts if it carries this run's nonce, which rules out stale traffic and lucky guesses. The fixtures also return a marker derived from the nonce that only the server knows. When that marker appears in the client's output, it shows the server's response actually made it back to the model.

What each status means

Supported
Every expected request was observed and the marker reached the model.
Partial
Some expected steps happened and some did not. A client that lists tools but never calls one lands here, as does one that calls a tool but drops the result.
Unsupported
The client connected to the server and finished its run, but never made any of the expected requests.
Test failed
The test could not be judged: the client errored, timed out, or never connected. This says nothing about whether the feature is supported.
Not tested
There is no result. This is never stored as a value and never shown as unsupported; an empty cell only means nobody has checked yet.

Protocol versions

MCP now has two generations in the field. Revisions up to 2025-11-25 open with an initialize handshake, and revision 2026-07-28 drops the handshake so every request carries its own version. The test server accepts both, and the version shown for each result is read from the client's actual requests. A client's documentation can say one thing while a particular install, configuration, or runtime does another, so the observed version is reported per test.

Reproducibility and history

Every test page shows the case definition that ran, a hash of that definition, the lab version, the client version, and the recorded traffic. Results are append only: a new run adds a new result and the old ones stay, so a regression shows up as history rather than a silent change. The full data set, including that traffic, is published as compat.json.

Some things are out of scope for now: HTTP transports, OAuth flows such as dynamic client registration and client ID metadata documents, and features that exist only in the newest revision. They will appear as new tests rather than as guesses.

Back to the results

© 2026 ABWaters. Thinking out loud.