Compare a run with a baseline
--baseline compares a new feedback bundle with an earlier run in the same results database. Capture the ID produced by your own first run; IDs printed in documentation do not exist in your database.
Run these commands from the project in Your first MCP test.
Create the baseline
Run the test once:
m3 test -- tests/test_m3_starter.pyThe command prints M3 run <id> and the local report path. Copy the value after M3 run into your shell.
For Bash or zsh:
printf 'Baseline run ID: '
IFS= read -r BASELINE_RUN_IDFor PowerShell:
$BASELINE_RUN_ID = Read-Host "Baseline run ID"Now make the behavior change you want to evaluate, then compare the next run.
m3 test --baseline "$BASELINE_RUN_ID" -- tests/test_m3_starter.pyPowerShell uses $BASELINE_RUN_ID without quotes around the variable name:
m3 test --baseline $BASELINE_RUN_ID -- tests/test_m3_starter.pyThe second command prints a different run ID. Capture it as CURRENT_RUN_ID the same way, then inspect .m3/reports/<CURRENT_RUN_ID>/feedback.json. The comparison reads the baseline; it does not modify it.
If M3 reports that the baseline was not found, confirm that both commands used the same project root and results database.