M3
Test MCP servers and the agents that use them.
M3 turns MCP interactions into repeatable Python tests. Run your existing pytest suite, capture tool calls and traces, compare runs, and explore results in a local browser viewer.
Install the CLI
With uv:
uv tool install sf-m3-cliOr use the shell installer on macOS or Linux:
curl -fsSL https://raw.githubusercontent.com/sineframe/m3/main/scripts/install-latest.sh | shGet started
From the project you want to test:
cd your-project
m3 init
m3 setup
m3 doctorm3 init creates a starter test. Replace its placeholder with a real assertion, then run m3 test -- tests/test_m3_starter.py. m3 setup installs the matching Python SDK into the project's environment.
What you can do
- Verify an MCP server's tools, schemas, responses, and error handling.
- Test direct MCP clients as well as agent sessions and harnesses.
- Exercise native Claude Code, OpenCode, Codex, and Pi harnesses, or connect another agent through an ACP-compatible adapter.
- Capture lifecycle events, tool calls, traces, artifacts, and evaluations.
- Repeat trials across harnesses and inspect aggregate results.
- Save executions to SQLite, produce deterministic feedback bundles, and compare later runs with
--baseline.
Test the behavior that matters
Mark one ordinary pytest test and let the CLI supply each harness and model:
import pytest
from m3 import expect
@pytest.mark.m3(suite_name="shipping", servers=[{
"type": "http", "url": "https://shipping.example.com/mcp", "trust": "public",
}])
def test_shipping(agent, server):
result = agent.run(
"Get a local shipping quote",
server=server,
permission_policy="allow",
)
expect(result).to_have_tool_call("shipping_quote")Replace the URL with your MCP endpoint. M3 supplies agent and server.
Set credentials with exported OPENCODE_API_KEY and OPENAI_API_KEY, or use an explicitly requested .env file. Then run two selections for two trials:
m3 test --env-file .env \
--harness opencode=opencode/big-pickle \
--harness codex=gpt-5.6-sol \
--trials 2 \
-- tests/test_shipping.pyThis collects four agent items and performs four executions.
To test specific harness releases, select a managed runtime and put each version after the harness name:
m3 test --runtime=managed \
--harness [email protected]=opencode/big-pickle \
--harness [email protected]=opencode/big-pickle \
-- tests/test_shipping.pyM3 downloads each release into a per-user cache, runs each selection with its own executable, and records the resolved harness and model in results and traces. See the CLI guide.
Bring your own harness
Built-in harnesses are convenient, but you can bring any agent implementing Agent Client Protocol (ACP) v1. Provide an agent dictionary with a manifest describing the executable, arguments, protocol version, and environment-variable references:
{
"schema_version": "m3.harness.v1",
"protocol": "acp",
"protocol_version": 1,
"command": "your-agent",
"args": ["--acp"],
"env": {"MY_AGENT_API_KEY": "${MY_AGENT_API_KEY}"}
}Bind the MCP server using a transport the agent advertises and supports; Streamable HTTP and stdio are available where applicable. M3 validates the manifest, can check local readiness and probe the configured process, then records the agent turn and captured MCP tool evidence using the same assertions as native harnesses.
See the BYO ACP guide and complete ACP example for the manifest shape, probes, and executable test.
Ask your coding agent to get started
If you use a coding agent, M3 includes a reusable testing-with-m3 skill. Ask your preferred agent to install the skill from this repository and use it to create tests for your server or agent workflow.
The skill helps an agent inspect the real MCP contract, choose direct server tests or agent-behavior tests, assert captured tool evidence, and iterate using feedback reports and baselines. The recommended workflow is to use the CLI to run the tests and inspect persistent history or the local UI.
Install and use the M3 skill from
https://github.com/sineframe/m3/tree/686c5f822e3abc950c0d80f942c62b127756637a/skills/testing-with-m3
Read its testing patterns, then add and run the smallest tests that verify
<the behavior I care about> against <my MCP server or agent workflow>.
If this project has no M3 test yet, run m3 init first and replace
its skipped starter test after inspecting the real server contract.
Keep direct server checks separate from agent tool-selection checks, and use
the M3 CLI to run the tests and inspect the resulting report or UI.This is one onboarding workflow; M3 works with any agent and any ordinary Python and pytest workflow.
Start with the CLI
The standalone m3 command runs your existing pytest suite in its project environment and includes the local browser viewer. No Node.js or frontend checkout is needed in the project under test. The CLI guide covers installation, project setup, environment checks, test selection, persistent results, UI use, and troubleshooting.
For CI, run m3 ci test; add --upload to publish the completed report. See the CI and credentials guide for secrets, GitHub Actions, authentication, and report retries.
Start with the CLI installation guide, then follow the project setup and testing guide.
From the project root, run m3 init to answer the project and suite name questions and create a skipped starter test and .env.example key template. Then run m3 setup, fill in the test, and use m3 test --env-file .env when the test needs provider keys.
CLI-managed runs use .m3/executions.sqlite by default and write an agent-readable report to .m3/reports/<run-id>/feedback.json. The CLI guide explains how to select pytest arguments, compare a run with a baseline, and open the bundled viewer. From that same project directory, m3 ui opens saved test runs without starting another test; m3 test --ui runs pytest first and then opens the viewer.
Use the SDK with direct pytest
The SDK is the library layer for Python tests. When managed history and the viewer are not needed, users may run SDK tests directly with pytest. The SDK README and quick start cover installation, direct clients, agent sessions, typed assertions, async tests, and explicit persistence. The CLI itself also runs pytest in the project's environment; these are two ways to run the same test style.
Choose your next step
| I want to… | Read |
|---|---|
| Install and run the standalone command | CLI guide |
| Write direct Python or pytest tests | SDK quick start |
| Understand traces, storage, and evaluations | SDK concepts |
| Score repeated agent trials | Evaluation guide |
| Test a deployed Streamable HTTP server | HTTP guide |
| Give a coding agent M3 instructions | Testing skill |
| Contribute to the implementation | Architecture |
Development
Contributors should start with the architecture guide, then use the package-specific guides for SDK, app, and CLI workflows. See the release guide when preparing a release.