---
title: "m3"
description: "Public Python API reference for m3."
---

> M3 v0.2.21 · commit 5a8005b4cfa8b984fafb0992dfc5d5cc4afa58a1.


# `m3`

Signatures use `...` for factory-backed or opaque defaults. Model field
tables show required status, defaults, constraints, and descriptions.

## `__version__`

`m3.__version__`

Installed distribution version.

## `MCPTestKit`

```python
m3.MCPTestKit(
    config: Config | _Mapping[str, _Any] | None = None,
    *,
    env: _Mapping[str, str] | None = None,
    cwd: str | _Path | None = None,
    probe_timeout_seconds: float = 5.0,
    probe_output_limit: int = 65536,
    store: _ExecutionStore | None = None,
    embedded_worker: bool = True,
    adapter_registry: _HarnessAdapterRegistry | None = None,
    run_id: _RunId | str | None = None,
    suite_name: str | None = None,
    project_id: _ProjectId | str | None = None,
    record_checks: bool = False,
    max_judge_requests: int | None = None,
    harness_cache_dir: str | _Path | None = None,
) -> None
```

Lifecycle-safe synchronous configuration and capability shell.

- `probes` (property): Synchronous capability namespace owned by this kit.
- `store` (property): The optional execution store configured on this kit.
- `run_id` (property)

```python
get_trace(
    self,
    execution_id: _ExecutionId | str,
) -> _TraceResult
```
Return the finalized stable trace for an execution.

```python
get_trace_view(
    self,
    execution_id: _ExecutionId | str,
) -> TraceView
```
Return the finalized typed trace view for an execution.

```python
read_raw_evidence(
    self,
    reference: _EvidenceRef,
    *,
    max_bytes: int = 1048576,
) -> RawEvidence
```
Read bounded, redacted raw evidence by its durable reference.

```python
close(
    self,
) -> None
```
Close the shell; repeated calls are intentionally harmless.

```python
capabilities(
    self,
    requests: _Iterable[ProbeRequest] = (),
) -> ProbeReport
```
Return the baseline or exactly the explicitly requested probes.

```python
register_evaluator(
    self,
    name: str,
    evaluator: _EvaluatorCallable,
) -> None
```
Register an evaluator callback by its serializable name.

```python
evaluate(
    self,
    subject: _Any,
    evaluator: str | _EvaluatorCallable,
    *,
    required: bool = False,
    goal: str | None = None,
    trace: _Any = None,
    artifacts: _Any = (),
    metadata: _Mapping[str, str | int | float | bool | None] | None = None,
    execution_id: _Any = None,
    turn_id: _Any = None,
    case_id: str | None = None,
) -> _EvaluationResult
```
Run and persist one evaluation without changing lifecycle.

```python
judge_response(
    self,
    *,
    name: str,
    input: str,
    actual: str,
    expected: str,
    judge: _LLMJudge,
    required: bool = False,
    execution_id: _Any = None,
    turn_id: _Any = None,
    case_id: str | None = None,
) -> _EvaluationResult
```
Judge one response and persist the result through this kit's runner.

```python
evaluation_results(
    self,
) -> tuple[_EvaluationResult, ...]
```

```python
agents(
    self,
    selections: _Any,
    *,
    trials: int = 1,
) -> tuple[_Any, ...]
```
Expand ordered agent dictionaries without starting any I/O.

```python
run(
    self,
    spec: _DirectSpec | _AgentSpec,
) -> _ExecutionResult
```

```python
submit(
    self,
    spec: _ExecutionSpec,
    *,
    human_input: _HumanInput = 'fail',
) -> ExecutionHandle
```

```python
direct(
    self,
    server: _ServerValue | _ServerBinding,
    *,
    protocol: object | None = None,
    timeout: float | None = None,
    validate_schemas: bool = False,
    secret_resolver: _Any = None,
    bearer_token: _Any = None,
    auth: _Any = None,
    for_agent: bool = False,
    resolve_host: _Any = None,
    raise_server_exceptions: bool = True,
    sampling_callback: _Any = None,
    elicitation_callback: _RemovedElicitationCallback = ...,
    list_roots_callback: _Any = None,
    logging_callback: _Any = None,
    message_handler: _Any = None,
    client_info: _Any = None,
    log_level: _Any = None,
    sampling_capabilities: _Any = None,
    result_claims: _Any = None,
    extensions: _Mapping[str, _Mapping[str, _Any]] | None = None,
    notification_bindings: _Iterable[_Any] | None = None,
    dispatcher: _Any = None,
    trace_bridge: _Any = None,
    trace_owner: bool = True,
    workspace_root: str | None = None,
) -> DirectClient
```

```python
agent_session(
    self,
    spec: _AgentSpec,
    *,
    adapter: _AgentAdapter | None = None,
    runtime_servers: _Iterable[_Any] = (),
    interaction_handlers: InteractionHandlers | None = None,
    harness_cache_dir: str | _Path | None = None,
) -> AgentSession
```

## `StdioServer`

```python
m3.StdioServer(
    *,
    name: str,
    trust: m3.types.TrustLevel = TrustLevel.UNTRUSTED,
    kind: Literal['stdio'] = 'stdio',
    command: str,
    args: tuple[str, ...] = (),
    environment: collections.abc.Mapping[str, m3.types.SecretReference | str] = ...,
    cwd: str | None = None,
) -> None
```

Model fields:

| Field | Type | Required | Default | Constraints | Description |
| --- | --- | --- | --- | --- | --- |
| `name` | `str` | Yes | — | `min_length=1, max_length=256` | — |
| `trust` | `m3.types.TrustLevel` | No | `TrustLevel.UNTRUSTED ('untrusted')` | — | — |
| `kind` | `Literal['stdio']` | No | `'stdio'` | — | — |
| `command` | `str` | Yes | — | `min_length=1` | — |
| `args` | `tuple[str, ...]` | No | `()` | — | — |
| `environment` | `collections.abc.Mapping[str, m3.types.SecretReference \| str]` | No | `factory builtins.dict()` | — | — |
| `cwd` | `str \| None` | No | `None` | — | — |

## `HTTPServer`

```python
m3.HTTPServer(
    *,
    name: str,
    trust: m3.types.TrustLevel = TrustLevel.UNTRUSTED,
    kind: Literal['streamable_http'] = 'streamable_http',
    url: str,
    headers: collections.abc.Mapping[str, m3.types.SecretReference | str] = ...,
    loopback_only: bool = False,
) -> None
```

Model fields:

| Field | Type | Required | Default | Constraints | Description |
| --- | --- | --- | --- | --- | --- |
| `name` | `str` | Yes | — | `min_length=1, max_length=256` | — |
| `trust` | `m3.types.TrustLevel` | No | `TrustLevel.UNTRUSTED ('untrusted')` | — | — |
| `kind` | `Literal['streamable_http']` | No | `'streamable_http'` | — | — |
| `url` | `str` | Yes | — | `min_length=1` | — |
| `headers` | `collections.abc.Mapping[str, m3.types.SecretReference \| str]` | No | `factory builtins.dict()` | — | — |
| `loopback_only` | `bool` | No | `False` | — | — |

## `InProcessServer`

```python
m3.InProcessServer(
    *,
    name: str,
    trust: m3.types.TrustLevel = TrustLevel.SDK_LOOPBACK,
    kind: Literal['in_process'] = 'in_process',
    factory: Any,
    descriptor: collections.abc.Mapping[str, Any] = ...,
    origin: str = 'python_registration',
) -> None
```

Runtime-only server descriptor; the factory is excluded from serialization.

Model fields:

| Field | Type | Required | Default | Constraints | Description |
| --- | --- | --- | --- | --- | --- |
| `name` | `str` | Yes | — | `min_length=1, max_length=256` | — |
| `trust` | `m3.types.TrustLevel` | No | `TrustLevel.SDK_LOOPBACK ('sdk_loopback')` | — | — |
| `kind` | `Literal['in_process']` | No | `'in_process'` | — | — |
| `factory` | `Any` | Yes | — | — | — |
| `descriptor` | `collections.abc.Mapping[str, Any]` | No | `factory builtins.dict()` | — | — |
| `origin` | `str` | No | `'python_registration'` | `min_length=1, max_length=256` | — |

## `ExecutionResult`

```python
m3.ExecutionResult(
    *,
    snapshot: m3.types.ExecutionState,
    turns: tuple[m3.types.TurnResult, ...] = (),
    trace: m3.types.TraceResult | None = None,
    direct_result: m3.types.ListToolsResult | m3.types.ListResourcesResult | m3.types.ListTemplatesResult | m3.types.ListPromptsResult | m3.types.CallToolResult | m3.types.ReadResourceResult | m3.types.GetPromptResult | m3.types.PingResult | None = None,
    evaluations: tuple[m3.types.EvaluationResult, ...] = (),
    artifacts: tuple[m3.types.ArtifactRef, ...] = (),
    activity_health: m3.types.ActivityHealth = ActivityHealth.NO_CALLS,
    error: m3.types.ErrorInfo | None = None,
    provenance: m3.types.SessionSource | None = None,
) -> None
```

Model fields:

| Field | Type | Required | Default | Constraints | Description |
| --- | --- | --- | --- | --- | --- |
| `snapshot` | `m3.types.ExecutionState` | Yes | — | — | — |
| `turns` | `tuple[m3.types.TurnResult, ...]` | No | `()` | — | — |
| `trace` | `m3.types.TraceResult \| None` | No | `None` | — | — |
| `direct_result` | `m3.types.ListToolsResult \| m3.types.ListResourcesResult \| m3.types.ListTemplatesResult \| m3.types.ListPromptsResult \| m3.types.CallToolResult \| m3.types.ReadResourceResult \| m3.types.GetPromptResult \| m3.types.PingResult \| None` | No | `None` | `variant 1: discriminator='kind'` | — |
| `evaluations` | `tuple[m3.types.EvaluationResult, ...]` | No | `()` | — | — |
| `artifacts` | `tuple[m3.types.ArtifactRef, ...]` | No | `()` | — | — |
| `activity_health` | `m3.types.ActivityHealth` | No | `ActivityHealth.NO_CALLS ('no_calls')` | — | — |
| `error` | `m3.types.ErrorInfo \| None` | No | `None` | — | — |
| `provenance` | `m3.types.SessionSource \| None` | No | `None` | — | — |
- `trace_view` (property): Return the finalized typed view for this execution trace.

## `ExecutionOutcome`

```python
m3.ExecutionOutcome(
    *values,
)
```

- `COMPLETED` = `'completed'`
- `FAILED` = `'failed'`
- `TIMED_OUT` = `'timed_out'`
- `CANCELLED` = `'cancelled'`
- `INTERRUPTED` = `'interrupted'`

## `TurnOutcome`

```python
m3.TurnOutcome(
    *values,
)
```

- `COMPLETED` = `'completed'`
- `FAILED` = `'failed'`
- `TIMED_OUT` = `'timed_out'`
- `CANCELLED` = `'cancelled'`
- `INTERRUPTED` = `'interrupted'`

## `expect`

```python
m3.expect(
    subject: _SubjectT,
    *,
    redaction_config: _RedactionConfig | None = None,
) -> Expectation[_SubjectT]
```

## `check`

```python
m3.check(
    *,
    redaction_config: _RedactionConfig | None = None,
) -> CheckGroup
```

## `MCPError`

```python
m3.MCPError(
    message: str,
    *,
    details: _Mapping[str, _Any] | None = None,
) -> None
```

Base class for expected M3 failures.
