Skip to content
ReferenceAlso known as: M3 execution timeout, capture_incomplete, trace limitations

CLI output and trace limitations ​

Lines printed after a run ​

LineMeaningWhat to do
M3 runThe run ID created by this invocation.Use it with --baseline and m3 upload.
M3 feedbackPath of this run's feedback.json.Read selected fields from this path.
M3 verdictsCount of test cases per verdict.Informational.
M3 observationsTool-error results and completed executions in this run.A tool error is a server result with is_error=True, not a test failure by itself.
M3: no tests executed; skipped-only runs failEvery selected test was skipped or deselected, and the run fails.Unskip or select a test.
M3 execution timeoutOne execution reached a deadline. It is printed once per timed-out execution: id, stage, elapsed, feedback.Informational. The run's result is the pytest outcome, so a test that expects a timeout can still pass. Inspect the execution's diagnostics when the timeout was not expected.
M3 test manifest persistence was incompletePytest results were not fully saved.Saved history for this run is incomplete; do not use it as a baseline.
M3 required evaluations blocked finalizationA required=True evaluation did not pass.The run is not successful even if test code caught the exception.
M3 CI excludedm3 ci test only: the number of tests excluded by markers with ci=False.Informational.
M3 feedback export failedfeedback.json was not written.The run's database records remain; rerun to get a feedback file.

Trace limitations ​

A trace's limitations list says which evidence M3 could not fully observe or save. Check it before treating a missing value as evidence of absence.

CodeWhere it appearsRoutine?
capture_incompleteDirect traces: always present. Every finalized direct trace records it, with completeness partial (sdk/src/m3/direct_trace.py:604-605). Agent traces: added when harness output was invalid, truncated, timed out, or duplicated.Routine on direct traces. On agent traces, read the trace diagnostics before concluding that a call did not happen.
capture_disabledProvider message and reasoning observations are dropped when provider-message capture is disabled (harness/observation_sink.py:85-89). Raw evidence is also not stored when raw capture is disabled (:123-126).Routine when the corresponding capture is off. Other normalized harness observations are still recorded.
partial_traceAn ACP turn failed after recording only part of its evidence (harness/acp.py:1730).No. The turn's evidence is incomplete.
cleanup_failedA server process or harness could not be shut down cleanly.No. Results stand, but check for leftover processes.
persistence_failedSome observations could not be saved (harness/observation_sink.py:120,403).No. The saved trace lacks evidence; do not treat a missing value as absence.