Two different questions: server contract and agent behavior
Testing an MCP server means answering two separate questions. The first is whether the server keeps its contract: it advertises the tools you expect, validates inputs, returns the structured results you promise, and reports errors the way clients can handle. The second is whether an agent that connects to the server actually uses it well: whether it calls the right tool, with the right arguments, in the right order.
M3 keeps those two questions apart on purpose. A direct test controls the MCP operation itself: it chooses a server, a method, arguments, and the expected result, and it needs no model. An agent test gives an agent a goal and checks the interaction M3 observed: which tool was called, with what arguments, and what the server returned.
An agent’s final answer is not evidence that it called the expected tool, so M3 asserts on recorded tool evidence instead. You assert on the final response separately, and only when that response is part of the contract.
Install M3 and run a first test
M3 needs Python 3.10 or newer and uv. The CLI and the project SDK are installed separately: the standalone CLI provides the m3 command and the local results viewer, and m3 setup installs the matching SDK, with pytest and storage support, into your project environment. Keep the two on matching releases.
The first-test walkthrough in the docs starts a small local server, discovers its tool, calls it, checks the structured response, then lets you break the assertion on purpose and open the saved run. It needs no model credentials, M3 account, or external server.
m3 init creates a skipped pytest starter. Replace it with a real assertion before relying on results: a run in which every selected test is skipped or deselected fails.
uv tool install sf-m3-cli
mkdir shipping-test && cd shipping-test
m3 init
m3 setup
m3 doctor
m3 test -- tests/test_m3_starter.py
m3 uiTest the server directly over stdio or HTTP
Direct tests call the MCP protocol without a model. They are the right place for tool schemas, structured results, protocol errors, resources, and prompts. Use a StdioServer when M3 should start your local server process for each test connection, and an HTTPServer when the Streamable HTTP endpoint is already running. M3 does not start or stop a deployed HTTP service, and HTTPServer defaults to untrusted, so a loopback test sets TRUSTED_PRIVATE explicitly. For authentication, use a header or secret reference and never put a credential in the URL.
list_all_tools() follows pagination until the server has no next page, so you can assert on the complete tool list. With validate_schemas=True, the client rejects arguments that do not match a tool’s advertised input schema before sending the call. An expected application error is a normal result with is_error set, which you assert on like any other value.
Dependent calls stay inside one direct-client context, which keeps one MCP connection open. That is how you test state across tool calls: create an order, capture the returned identifier, and retrieve it in a later call. M3 starts a fresh server process for each test, so state does not leak between tests. The docs also cover resources, prompts, and observing a tool-list change on a subscription stream.
A passing result means M3 connected, discovered the tool, and observed the structured result you asserted. It does not certify behavior outside the inputs you asserted, so assert the value that matters rather than only that the call did not error.
import sys
from pathlib import Path
import pytest
from m3 import MCPTestKit, StdioServer
pytestmark = pytest.mark.m3(suite_name="shipping")
def test_shipping_quote() -> None:
"""shipping_quote prices a 2 kg local parcel at 9.00 USD."""
project_root = Path(__file__).parents[1]
server = StdioServer(
name="shipping",
command=sys.executable,
args=(str(project_root / "shipping_server.py"),),
cwd=str(project_root),
)
with MCPTestKit(env={}) as kit, kit.direct(server) as client:
tools = client.list_all_tools()
assert [tool.name for tool in tools] == ["shipping_quote"]
result = client.call_tool("shipping_quote", {"weight_kg": 2, "zone": "local"})
assert result.is_error is False
assert result.structured_content == {"amount": 9.0, "currency": "USD"}Test agent tool selection with real harness binaries
Agent tests add a harness and a model so you can inspect the tool calls the agent actually made. M3 connects to Claude Code, Codex, OpenCode, and Pi through native adapters, and to any ACP-compatible agent through an ACP v1 adapter, where you install and manage the executable. Each harness needs its installed executable and provider configuration, and these adapters do not have identical capabilities, so check the compatibility reference for the evidence and approval behavior your test needs.
The test limits the agent to the tools you name, then asserts on the recorded call. In the docs example the selection is restricted to shipping:shipping_quote, and the matcher checks the tool name, server alias, arguments, a successful result, and the call count. If the agent completes without making that call, the assertion fails even when its reply claims it used the tool.
To compare agents, keep the server, prompt, and assertion fixed and vary the harness. Run the same cases as a matrix across harnesses and trials, or pin native harness versions with m3 test --runtime managed so each execution records the requested and resolved runtime. Each trial is an independent execution rather than a retry, and a matrix expands test work without making model behavior deterministic.
Credentials stay out of test code. A credential mapping such as OPENAI_API_KEY=MY_OPENAI_KEY lists variable names, not secret values, and each native adapter isolates its child environment. Agent runs depend on the harness, model, provider configuration, and approval behavior, so record those inputs when you compare results, and expect agent tests to need provider access and to incur cost.
expect(result).to_have_tool_call(
"shipping_quote",
turn=turn,
server="shipping",
arguments={"weight_kg": 2, "zone": "local"},
status="success",
count=1,
)Keep the evidence: traces, saved runs, baselines, and the viewer
M3 records what it can observe at the transport and harness boundaries. After a client or agent session closes, the finalized trace holds the complete observed execution: calls, results, messages, timing, and outcome. Trace values distinguish an observed value from an unavailable one, and each trace lists its limitations so you do not mistake missing data for proof of absence.
m3 test saves run history to a local SQLite database under .m3/ in your project. m3 ui opens that history in a local browser viewer bound to loopback with a launch token, and m3 test --ui opens the run you just completed. Treat the printed viewer URL as a credential while it is running.
To check a change against an earlier run, capture the run ID that m3 test prints and pass it back with --baseline. The resulting feedback compares tool interfaces, tests, failures, and evaluations, and it marks comparisons as unlike-for-like when a judge model or rubric changed. A fresh worktree or CI job has no history, because .m3/ is ignored, so preserve a database if you want baselines there.
When you are ready to run the same suite in automation, see the MCP CI/CD guide. For assertions beyond a single call, M3 also supports custom evaluators and LLM judges, covered in the evaluations guides.
Frequently asked questions
- How do I test an MCP server?
- Install the M3 CLI with uv tool install sf-m3-cli, run m3 init and m3 setup in your project, and write a pytest test that connects to your server with a StdioServer or HTTPServer, lists its tools, calls one with known arguments, and asserts on the structured result. Run it with m3 test.
- Do I need a model or API key to test an MCP server?
- Not for direct tests. They call the MCP protocol without a model and need no provider key or M3 account. Agent tests need a native harness such as Claude Code, Codex, OpenCode, or Pi, or an ACP-compatible agent, plus provider access, and they may incur cost.
- Can M3 test both stdio and HTTP MCP servers?
- Yes. Use StdioServer when M3 should start a local server process for each test, and HTTPServer for a running Streamable HTTP endpoint. M3 does not start or stop a deployed HTTP service, and HTTPServer defaults to untrusted unless you set a trust level.
- Which agents can M3 test an MCP server with?
- Claude Code, Codex, OpenCode, and Pi through native adapters, and any ACP-compatible agent through the ACP v1 adapter. Harness availability does not imply identical evidence, approval, or cancellation behavior, so check the compatibility reference for the capability you need.
- Does M3 work with pytest?
- Yes. Tests are ordinary pytest tests. m3 test runs pytest in your project environment and saves run history, and options after -- are passed to pytest unchanged. The pytest plugin adds M3 fixtures, selection, storage integration, and feedback generation.
- Is M3 ready for production use?
- M3 is a pre-release preview, so features may change. Pin the CLI and project SDK to matching releases, and check the compatibility reference for current boundaries.