Skip to content
Reference

m3 ​

Signatures use ... for factory-backed or opaque defaults. Model field tables show required status, defaults, constraints, and descriptions.

__version__ ​

m3.__version__

Installed distribution version.

MCPTestKit ​

python
m3.MCPTestKit(
    config: Config | _Mapping[str, _Any] | None = None,
    *,
    env: _Mapping[str, str] | None = None,
    cwd: str | _Path | None = None,
    probe_timeout_seconds: float = 5.0,
    probe_output_limit: int = 65536,
    store: _ExecutionStore | None = None,
    embedded_worker: bool = True,
    adapter_registry: _HarnessAdapterRegistry | None = None,
    run_id: _RunId | str | None = None,
    suite_name: str | None = None,
    project_id: _ProjectId | str | None = None,
    record_checks: bool = False,
    max_judge_requests: int | None = None,
    harness_cache_dir: str | _Path | None = None,
) -> None

Lifecycle-safe synchronous configuration and capability shell.

  • probes (property): Synchronous capability namespace owned by this kit.
  • store (property): The optional execution store configured on this kit.
  • run_id (property)
python
get_trace(
    self,
    execution_id: _ExecutionId | str,
) -> _TraceResult

Return the finalized stable trace for an execution.

python
get_trace_view(
    self,
    execution_id: _ExecutionId | str,
) -> TraceView

Return the finalized typed trace view for an execution.

python
read_raw_evidence(
    self,
    reference: _EvidenceRef,
    *,
    max_bytes: int = 1048576,
) -> RawEvidence

Read bounded, redacted raw evidence by its durable reference.

python
close(
    self,
) -> None

Close the shell; repeated calls are intentionally harmless.

python
capabilities(
    self,
    requests: _Iterable[ProbeRequest] = (),
) -> ProbeReport

Return the baseline or exactly the explicitly requested probes.

python
register_evaluator(
    self,
    name: str,
    evaluator: _EvaluatorCallable,
) -> None

Register an evaluator callback by its serializable name.

python
evaluate(
    self,
    subject: _Any,
    evaluator: str | _EvaluatorCallable,
    *,
    required: bool = False,
    goal: str | None = None,
    trace: _Any = None,
    artifacts: _Any = (),
    metadata: _Mapping[str, str | int | float | bool | None] | None = None,
    execution_id: _Any = None,
    turn_id: _Any = None,
    case_id: str | None = None,
) -> _EvaluationResult

Run and persist one evaluation without changing lifecycle.

python
judge_response(
    self,
    *,
    name: str,
    input: str,
    actual: str,
    expected: str,
    judge: _LLMJudge,
    required: bool = False,
    execution_id: _Any = None,
    turn_id: _Any = None,
    case_id: str | None = None,
) -> _EvaluationResult

Judge one response and persist the result through this kit's runner.

python
evaluation_results(
    self,
) -> tuple[_EvaluationResult, ...]
python
agents(
    self,
    selections: _Any,
    *,
    trials: int = 1,
) -> tuple[_Any, ...]

Expand ordered agent dictionaries without starting any I/O.

python
run(
    self,
    spec: _DirectSpec | _AgentSpec,
) -> _ExecutionResult
python
submit(
    self,
    spec: _ExecutionSpec,
    *,
    human_input: _HumanInput = 'fail',
) -> ExecutionHandle
python
direct(
    self,
    server: _ServerValue | _ServerBinding,
    *,
    protocol: object | None = None,
    timeout: float | None = None,
    validate_schemas: bool = False,
    secret_resolver: _Any = None,
    bearer_token: _Any = None,
    auth: _Any = None,
    for_agent: bool = False,
    resolve_host: _Any = None,
    raise_server_exceptions: bool = True,
    sampling_callback: _Any = None,
    elicitation_callback: _RemovedElicitationCallback = ...,
    list_roots_callback: _Any = None,
    logging_callback: _Any = None,
    message_handler: _Any = None,
    client_info: _Any = None,
    log_level: _Any = None,
    sampling_capabilities: _Any = None,
    result_claims: _Any = None,
    extensions: _Mapping[str, _Mapping[str, _Any]] | None = None,
    notification_bindings: _Iterable[_Any] | None = None,
    dispatcher: _Any = None,
    trace_bridge: _Any = None,
    trace_owner: bool = True,
    workspace_root: str | None = None,
) -> DirectClient
python
agent_session(
    self,
    spec: _AgentSpec,
    *,
    adapter: _AgentAdapter | None = None,
    runtime_servers: _Iterable[_Any] = (),
    interaction_handlers: InteractionHandlers | None = None,
    harness_cache_dir: str | _Path | None = None,
) -> AgentSession

StdioServer ​

python
m3.StdioServer(
    *,
    name: str,
    trust: m3.types.TrustLevel = TrustLevel.UNTRUSTED,
    kind: Literal['stdio'] = 'stdio',
    command: str,
    args: tuple[str, ...] = (),
    environment: collections.abc.Mapping[str, m3.types.SecretReference | str] = ...,
    cwd: str | None = None,
) -> None

Model fields:

FieldTypeRequiredDefaultConstraintsDescription
namestrYes—min_length=1, max_length=256—
trustm3.types.TrustLevelNoTrustLevel.UNTRUSTED ('untrusted')——
kindLiteral['stdio']No'stdio'——
commandstrYes—min_length=1—
argstuple[str, ...]No()——
environmentcollections.abc.Mapping[str, m3.types.SecretReference | str]Nofactory builtins.dict()——
cwdstr | NoneNoNone——

HTTPServer ​

python
m3.HTTPServer(
    *,
    name: str,
    trust: m3.types.TrustLevel = TrustLevel.UNTRUSTED,
    kind: Literal['streamable_http'] = 'streamable_http',
    url: str,
    headers: collections.abc.Mapping[str, m3.types.SecretReference | str] = ...,
    loopback_only: bool = False,
) -> None

Model fields:

FieldTypeRequiredDefaultConstraintsDescription
namestrYes—min_length=1, max_length=256—
trustm3.types.TrustLevelNoTrustLevel.UNTRUSTED ('untrusted')——
kindLiteral['streamable_http']No'streamable_http'——
urlstrYes—min_length=1—
headerscollections.abc.Mapping[str, m3.types.SecretReference | str]Nofactory builtins.dict()——
loopback_onlyboolNoFalse——

InProcessServer ​

python
m3.InProcessServer(
    *,
    name: str,
    trust: m3.types.TrustLevel = TrustLevel.SDK_LOOPBACK,
    kind: Literal['in_process'] = 'in_process',
    factory: Any,
    descriptor: collections.abc.Mapping[str, Any] = ...,
    origin: str = 'python_registration',
) -> None

Runtime-only server descriptor; the factory is excluded from serialization.

Model fields:

FieldTypeRequiredDefaultConstraintsDescription
namestrYes—min_length=1, max_length=256—
trustm3.types.TrustLevelNoTrustLevel.SDK_LOOPBACK ('sdk_loopback')——
kindLiteral['in_process']No'in_process'——
factoryAnyYes———
descriptorcollections.abc.Mapping[str, Any]Nofactory builtins.dict()——
originstrNo'python_registration'min_length=1, max_length=256—

ExecutionResult ​

python
m3.ExecutionResult(
    *,
    snapshot: m3.types.ExecutionState,
    turns: tuple[m3.types.TurnResult, ...] = (),
    trace: m3.types.TraceResult | None = None,
    direct_result: m3.types.ListToolsResult | m3.types.ListResourcesResult | m3.types.ListTemplatesResult | m3.types.ListPromptsResult | m3.types.CallToolResult | m3.types.ReadResourceResult | m3.types.GetPromptResult | m3.types.PingResult | None = None,
    evaluations: tuple[m3.types.EvaluationResult, ...] = (),
    artifacts: tuple[m3.types.ArtifactRef, ...] = (),
    activity_health: m3.types.ActivityHealth = ActivityHealth.NO_CALLS,
    error: m3.types.ErrorInfo | None = None,
    provenance: m3.types.SessionSource | None = None,
) -> None

Model fields:

FieldTypeRequiredDefaultConstraintsDescription
snapshotm3.types.ExecutionStateYes———
turnstuple[m3.types.TurnResult, ...]No()——
tracem3.types.TraceResult | NoneNoNone——
direct_resultm3.types.ListToolsResult | m3.types.ListResourcesResult | m3.types.ListTemplatesResult | m3.types.ListPromptsResult | m3.types.CallToolResult | m3.types.ReadResourceResult | m3.types.GetPromptResult | m3.types.PingResult | NoneNoNonevariant 1: discriminator='kind'—
evaluationstuple[m3.types.EvaluationResult, ...]No()——
artifactstuple[m3.types.ArtifactRef, ...]No()——
activity_healthm3.types.ActivityHealthNoActivityHealth.NO_CALLS ('no_calls')——
errorm3.types.ErrorInfo | NoneNoNone——
provenancem3.types.SessionSource | NoneNoNone——
  • trace_view (property): Return the finalized typed view for this execution trace.

ExecutionOutcome ​

python
m3.ExecutionOutcome(
    *values,
)
  • COMPLETED = 'completed'
  • FAILED = 'failed'
  • TIMED_OUT = 'timed_out'
  • CANCELLED = 'cancelled'
  • INTERRUPTED = 'interrupted'

TurnOutcome ​

python
m3.TurnOutcome(
    *values,
)
  • COMPLETED = 'completed'
  • FAILED = 'failed'
  • TIMED_OUT = 'timed_out'
  • CANCELLED = 'cancelled'
  • INTERRUPTED = 'interrupted'

expect ​

python
m3.expect(
    subject: _SubjectT,
    *,
    redaction_config: _RedactionConfig | None = None,
) -> Expectation[_SubjectT]

check ​

python
m3.check(
    *,
    redaction_config: _RedactionConfig | None = None,
) -> CheckGroup

MCPError ​

python
m3.MCPError(
    message: str,
    *,
    details: _Mapping[str, _Any] | None = None,
) -> None

Base class for expected M3 failures.