Skip to content

M3 Python SDK ​

Use these pages to start writing tests for MCP servers and agent harnesses. Tests are ordinary executable pytest tests; run them through the M3 CLI for persisted results and an optional local UI, or invoke pytest directly. Persisted pytest runs require a suite name on each selected test; direct Python/notebook use and pytest without persisted results do not. See the quick start for markers and selection behavior. Native harness choices include Claude Code, OpenCode, Codex App Server, and Pi RPC; ACP remains the bring-your-own protocol path. The project SDK and standalone CLI are separate installations; adding the SDK to a project does not install the m3 command or its bundled UI.

  • Streamable HTTP — test a deployed MCP endpoint directly, through an agent session, or with a harness matrix.
  • Quick start — install the SDK, write the first real MCP test, and run it with the CLI or pytest.
  • Concepts — understand server bindings, lifecycle, results, schemas, chained workflows, traces, and isolation.
  • Elicitation — write composable MRTR tests using real either/or, optional, same-round, and later-round examples.
  • Elicitation API reference — exact helpers, action boundaries, manual input, managed rounds, and trace fields.
  • Pi-to-Codex MRTR parity inventory — scenario-by-scenario test mapping, Codex evidence, and pending gaps.
  • Examples — choose the example matching a deployed endpoint, local stdio server, harness, ACP agent, or matrix workflow.
  • Evaluations — define evaluators, score repeated agent trials across harnesses, save results, and query pass-rate trends.

The copyable typed observability starting point is test_typed_trace_view.py; it is also linked from the quick start and concepts.

For harness-driven tool assertions, start with test_harness_trace_view.py.

The examples run against the deterministic stdio server in example_mcp_server.py. They use only public M3 SDK APIs, so the same tests work with either runner.