SineFrameM3CI gate for MCPs Back to home
Blog

Your MCP server passed. Did the agent call it?

Test an MCP server with plain pytest, break it on purpose, then run the same tests through real Claude Code and Codex. A walkthrough with M3.

Published
October 4, 2026
Reading time
4 min

An MCP server can be completely correct and still fail its users.

The tool returns the right answer. The schema is valid. Every unit test is green. Then a coding agent gets the user's question, decides it already knows the answer, and never calls your tool. Nothing errors and nothing fails. Somebody just gets a made-up number.

Ordinary tests can't catch that, because they never run the agent. M3 is a pytest-based test runner for MCP servers and the agents that use them. This post walks through what that looks like on a small example: a shipping-quote server, 64 tests, one deliberate regression, and a real Claude Code and Codex run.

The server

The example server has one tool, shipping_quote. It takes a weight and a zone and returns a structured quote:

rates = {"local": 2.0, "regional": 3.5, "international": 7.0}
quote = {"amount": round(5 + weight * rates[zone], 2), "currency": "USD"}

Invalid input (zero or negative weight, an unknown zone) comes back as an MCP tool error instead of a crash.

Setup

uv tool install sf-m3-cli
m3 init      # project identity, a starter test, and the agent skill
m3 setup     # installs the matching SDK and pytest support into the project
m3 doctor    # checks the environment

No API keys are needed for any of the direct tests below. M3 talks to your server over stdio or Streamable HTTP, without a model in between.

Plain pytest against the real server

A direct test starts the actual server process, lists its tools and calls one:

import sys
from pathlib import Path

import pytest
from m3 import MCPTestKit, StdioServer

pytestmark = pytest.mark.m3(suite_name="shipping")

SERVER = StdioServer(
    name="shipping",
    command=sys.executable,
    args=(str(Path(__file__).parents[1] / "shipping_server.py"),),
    cwd=str(Path(__file__).parents[1]),
)


def quote(arguments: dict):
    with MCPTestKit(env={}) as kit, kit.direct(SERVER) as client:
        assert [tool.name for tool in client.list_all_tools()] == ["shipping_quote"]
        return client.call_tool("shipping_quote", arguments)


def test_local_quote() -> None:
    """A 2 kg local parcel costs 9.00 USD."""
    result = quote({"weight_kg": 2, "zone": "local"})
    assert result.is_error is False
    assert result.structured_content == {"amount": 9.0, "currency": "USD"}

Because it's plain pytest, everything you already know applies. Parametrizing once gives 60 cases: 16 weights across 3 zones, plus 12 invalid inputs that must come back as tool errors.

@pytest.mark.parametrize("zone", list(RATES))
@pytest.mark.parametrize("weight_kg", [0.25, 0.5, 1, 1.5, 2, 3, 4, 5, 7.5, 10, 15, 20, 25, 40, 60, 80])
def test_quote_matrix(weight_kg: float, zone: str) -> None:
    result = quote({"weight_kg": weight_kg, "zone": zone})
    expected = round(5 + weight_kg * RATES[zone], 2)
    assert result.structured_content == {"amount": expected, "currency": "USD"}

Run the whole suite on four workers:

m3 test -n 4 -- tests

M3's own options (like -n) go before --, and everything after -- goes straight to pytest. In a terminal you get a live grid with one square per test, filling in as the workers finish. Here it ends with 64 green squares and 64 passed.

Now break it on purpose

Here is a one-line edit to the server that looks harmless: the local rate goes from 2.0 to 2.5.

- rates = {"local": 2.0, "regional": 3.5, "international": 7.0}
+ rates = {"local": 2.5, "regional": 3.5, "international": 7.0}

Run the same command again:

m3 test -n 4 -- tests

Red squares land across the grid as the run goes. The result is 17 failed, 47 passed: all 16 local cases in the matrix, plus the original test_local_quote. The failure tells you exactly what moved:

E         Differing items:
E         {'amount': 10.0} != {'amount': 9.0}

The test expected 9.0 and the server now returns 10.0. That's the most boring possible bug, and that's the point. A rate change like this is easy to merge by accident, and the suite catches it before it reaches main.

The same tests, through real agents

Direct tests prove the server is right. They don't prove an agent will use it. For that, M3 runs real coding-agent binaries (Claude Code, Codex, OpenCode and Pi, or any ACP agent) against your server and records every MCP tool call they make.

An agent test asks a question and asserts on the call that actually happened, not on what the agent says it did:

from m3 import ExecutionOutcome, expect


@pytest.mark.parametrize(
    ("prompt", "arguments"),
    [
        ("Use the shipping server to price a 2 kg local parcel.", {"weight_kg": 2, "zone": "local"}),
        ("Use the shipping server to quote a 5 kg international parcel.", {"weight_kg": 5, "zone": "international"}),
    ],
    ids=["local-2kg", "international-5kg"],
)
def test_agent_quotes_with_the_tool(agent, prompt: str, arguments: dict) -> None:
    """The agent answers from the shipping_quote tool, not from memory."""
    result = agent.run(prompt, server=SERVER, timeout=180)
    assert result.snapshot.outcome is ExecutionOutcome.COMPLETED, result.error
    expect(result).to_have_tool_call(
        "shipping_quote", server="shipping", arguments=arguments, status="success", count=1
    )

The harnesses are chosen on the command line, so the test file doesn't change:

m3 test -n 4 --runtime managed \
  --harness claude-code=claude-sonnet-5-5 \
  --harness codex=gpt-5.6-luna \
  -- agents

Credentials come from the usual places: Claude Code picks up ANTHROPIC_API_KEY, and Codex can reuse your existing Codex login. --runtime managed downloads and pins the harness binaries, so the run records exactly which versions it used (here, Claude Code 2.1.289 and Codex 0.160.0). Two prompts times two harnesses gives four agent runs.

On the run in the video, Claude Code passed both cases. Codex passed the local one and missed the international one: the matcher didn't find a call that satisfied the expectation. That's exactly the kind of difference you want to know about before your users find it. The same test, the same server, and a different agent behaves differently.

Every tool call, kept

Every run is saved, and m3 ui opens a local report. If you upload runs from CI, the same report works in the hosted viewer.

For an agent run, the report shows each case with its agent and model, the failed assertion with expected and observed values, and the full trace: the prompt, every tool call with its arguments and result, reasoning steps, and the final answer.

One run worth showing is a probe where the server rejects the agent's first attempt. Claude Code called strict_echo with zone='europe', and the server answered:

Tool error: Invalid zone 'europe'. Fix: call strict_echo again with zone set to exactly one of: local, regional, global.

The trace then shows the agent reasoning about that error, retrying with zone=regional, and reporting the confirmation code. When an agent test fails, this is where you find out why, without rerunning anything.

A CI gate that zero tests can't pass

The same suite runs in CI with m3 ci test:

m3 ci test; echo "exit $?"
# ... 17 failed, 47 passed
# exit 1

It also refuses the most dangerous kind of green build, a run where nothing executed:

m3 ci test -- -k nothing_selected; echo "exit $?"
# ... 64 deselected
# exit 5

A typo in a -k filter or a broken collection step can't quietly approve a pull request. The GitHub Actions steps are a few lines:

- run: m3 setup --python .venv/bin/python
- run: m3 ci test --python .venv/bin/python -- tests/ -q

The full workflow, including checkout and uv, is in Run M3 in GitHub Actions.

Try it

uv tool install sf-m3-cli
m3 init
m3 setup
m3 doctor

Then follow Write your first MCP test: it uses a local server and needs no model credentials. Or hand the setup to your coding agent: m3 init installs a testing-with-m3 skill, so Claude Code or Codex can write the first tests for your server.

M3 is open source under Apache-2.0 and currently in alpha.

  • Docs: m3.sineframe.com/docs
  • Code: github.com/sineframe/m3

If you build MCP servers, we'd love to know which agent your users run most, and what would stop you from putting a test like this in CI.

RK

Written byRishav KatochCo-founder, SineFrame

Read this post as MarkdownAll posts
© 2026 SineFrame M3
HomeMCP testingMCP CI gateM3 vs MCPJamBlogDocumentationPrivacy policyTerms of use