# Your MCP server passed. Did the agent call it?

<video controls muted playsinline preload="metadata" poster="/blog/did-the-agent-call-it.jpg" src="/blog/did-the-agent-call-it.mp4"></video>

An MCP server can be completely correct and still fail its users.

The tool returns the right answer. The schema is valid. Every unit test is green. Then a coding agent gets the user's question, decides it already knows the answer, and never calls your tool. Nothing errors and nothing fails. Somebody just gets a made-up number.

Ordinary tests can't catch that, because they never run the agent. **M3** is a pytest-based test runner for MCP servers *and* the agents that use them. This post walks through what that looks like on a small example: a shipping-quote server, 64 tests, one deliberate regression, and a real Claude Code and Codex run.

## The server

The example server has one tool, `shipping_quote`. It takes a weight and a zone and returns a structured quote:

```python
rates = {"local": 2.0, "regional": 3.5, "international": 7.0}
quote = {"amount": round(5 + weight * rates[zone], 2), "currency": "USD"}
```

Invalid input (zero or negative weight, an unknown zone) comes back as an MCP tool error instead of a crash.

## Setup

```sh
uv tool install sf-m3-cli
m3 init      # project identity, a starter test, and the agent skill
m3 setup     # installs the matching SDK and pytest support into the project
m3 doctor    # checks the environment
```

No API keys are needed for any of the direct tests below. M3 talks to your server over stdio or Streamable HTTP, without a model in between.

## Plain pytest against the real server

A direct test starts the actual server process, lists its tools and calls one:

```python
import sys
from pathlib import Path

import pytest
from m3 import MCPTestKit, StdioServer

pytestmark = pytest.mark.m3(suite_name="shipping")

SERVER = StdioServer(
    name="shipping",
    command=sys.executable,
    args=(str(Path(__file__).parents[1] / "shipping_server.py"),),
    cwd=str(Path(__file__).parents[1]),
)


def quote(arguments: dict):
    with MCPTestKit(env={}) as kit, kit.direct(SERVER) as client:
        assert [tool.name for tool in client.list_all_tools()] == ["shipping_quote"]
        return client.call_tool("shipping_quote", arguments)


def test_local_quote() -> None:
    """A 2 kg local parcel costs 9.00 USD."""
    result = quote({"weight_kg": 2, "zone": "local"})
    assert result.is_error is False
    assert result.structured_content == {"amount": 9.0, "currency": "USD"}
```

Because it's plain pytest, everything you already know applies. Parametrizing once gives 60 cases: 16 weights across 3 zones, plus 12 invalid inputs that must come back as tool errors.

```python
@pytest.mark.parametrize("zone", list(RATES))
@pytest.mark.parametrize("weight_kg", [0.25, 0.5, 1, 1.5, 2, 3, 4, 5, 7.5, 10, 15, 20, 25, 40, 60, 80])
def test_quote_matrix(weight_kg: float, zone: str) -> None:
    result = quote({"weight_kg": weight_kg, "zone": zone})
    expected = round(5 + weight_kg * RATES[zone], 2)
    assert result.structured_content == {"amount": expected, "currency": "USD"}
```

Run the whole suite on four workers:

```sh
m3 test -n 4 -- tests
```

M3's own options (like `-n`) go before `--`, and everything after `--` goes straight to pytest. In a terminal you get a live grid with one square per test, filling in as the workers finish. Here it ends with 64 green squares and `64 passed`.

## Now break it on purpose

Here is a one-line edit to the server that looks harmless: the local rate goes from 2.0 to 2.5.

```diff
- rates = {"local": 2.0, "regional": 3.5, "international": 7.0}
+ rates = {"local": 2.5, "regional": 3.5, "international": 7.0}
```

Run the same command again:

```sh
m3 test -n 4 -- tests
```

Red squares land across the grid as the run goes. The result is 17 failed, 47 passed: all 16 local cases in the matrix, plus the original `test_local_quote`. The failure tells you exactly what moved:

```text
E         Differing items:
E         {'amount': 10.0} != {'amount': 9.0}
```

The test expected 9.0 and the server now returns 10.0. That's the most boring possible bug, and that's the point. A rate change like this is easy to merge by accident, and the suite catches it before it reaches main.

## The same tests, through real agents

Direct tests prove the server is right. They don't prove an agent will use it. For that, M3 runs real coding-agent binaries (Claude Code, Codex, OpenCode and Pi, or any ACP agent) against your server and records every MCP tool call they make.

An agent test asks a question and asserts on the call that actually happened, not on what the agent says it did:

```python
from m3 import ExecutionOutcome, expect


@pytest.mark.parametrize(
    ("prompt", "arguments"),
    [
        ("Use the shipping server to price a 2 kg local parcel.", {"weight_kg": 2, "zone": "local"}),
        ("Use the shipping server to quote a 5 kg international parcel.", {"weight_kg": 5, "zone": "international"}),
    ],
    ids=["local-2kg", "international-5kg"],
)
def test_agent_quotes_with_the_tool(agent, prompt: str, arguments: dict) -> None:
    """The agent answers from the shipping_quote tool, not from memory."""
    result = agent.run(prompt, server=SERVER, timeout=180)
    assert result.snapshot.outcome is ExecutionOutcome.COMPLETED, result.error
    expect(result).to_have_tool_call(
        "shipping_quote", server="shipping", arguments=arguments, status="success", count=1
    )
```

The harnesses are chosen on the command line, so the test file doesn't change:

```sh
m3 test -n 4 --runtime managed \
  --harness claude-code=claude-sonnet-5-5 \
  --harness codex=gpt-5.6-luna \
  -- agents
```

Credentials come from the usual places: Claude Code picks up `ANTHROPIC_API_KEY`, and Codex can reuse your existing Codex login. `--runtime managed` downloads and pins the harness binaries, so the run records exactly which versions it used (here, Claude Code 2.1.289 and Codex 0.160.0). Two prompts times two harnesses gives four agent runs.

On the run in the video, Claude Code passed both cases. Codex passed the local one and missed the international one: the matcher didn't find a call that satisfied the expectation. That's exactly the kind of difference you want to know about before your users find it. The same test, the same server, and a different agent behaves differently.

## Every tool call, kept

Every run is saved, and `m3 ui` opens a local report. If you upload runs from CI, the same report works in the hosted viewer.

For an agent run, the report shows each case with its agent and model, the failed assertion with expected and observed values, and the full trace: the prompt, every tool call with its arguments and result, reasoning steps, and the final answer.

One run worth showing is a probe where the server rejects the agent's first attempt. Claude Code called `strict_echo` with `zone='europe'`, and the server answered:

```text
Tool error: Invalid zone 'europe'. Fix: call strict_echo again with zone set to exactly one of: local, regional, global.
```

The trace then shows the agent reasoning about that error, retrying with `zone=regional`, and reporting the confirmation code. When an agent test fails, this is where you find out why, without rerunning anything.

## A CI gate that zero tests can't pass

The same suite runs in CI with `m3 ci test`:

```sh
m3 ci test; echo "exit $?"
# ... 17 failed, 47 passed
# exit 1
```

It also refuses the most dangerous kind of green build, a run where nothing executed:

```sh
m3 ci test -- -k nothing_selected; echo "exit $?"
# ... 64 deselected
# exit 5
```

A typo in a `-k` filter or a broken collection step can't quietly approve a pull request. The GitHub Actions steps are a few lines:

```yaml
- run: m3 setup --python .venv/bin/python
- run: m3 ci test --python .venv/bin/python -- tests/ -q
```

The full workflow, including checkout and uv, is in [Run M3 in GitHub Actions](https://m3.sineframe.com/docs/guides/ci/github-actions).

## Try it

```sh
uv tool install sf-m3-cli
m3 init
m3 setup
m3 doctor
```

Then follow [Write your first MCP test](https://m3.sineframe.com/docs/getting-started): it uses a local server and needs no model credentials. Or hand the setup to your coding agent: `m3 init` installs a `testing-with-m3` skill, so Claude Code or Codex can write the first tests for your server.

M3 is open source under Apache-2.0 and currently in alpha.

- Docs: [m3.sineframe.com/docs](https://m3.sineframe.com/docs/)
- Code: [github.com/sineframe/m3](https://github.com/sineframe/m3)

If you build MCP servers, we'd love to know which agent your users run most, and what would stop you from putting a test like this in CI.
