SineFrameM3CI gate for MCPs Back to home
MCP CI/CD

MCP CI/CD: test your MCP server in GitHub Actions

A CI gate for an MCP server should fail when behavior regresses and should run the same checks you run on your machine. SineFrame M3 does that with m3 ci test: ordinary pytest tests, a failing exit on a failed run, and optional publishing of the result. M3 is a pre-release preview.

Facts last checked October 4, 2026

What the gate checks

M3 tests are Python and pytest tests, so a CI job runs them the same way a developer does. A test fails when its assertion fails. In the docs’ examples a failing test exits with code 1, and m3 ci test still writes the run’s report in that case, so you can attach it to the job.

A gate that can pass on nothing is not a gate. If every selected test is skipped or deselected, M3 prints "no tests executed; skipped-only runs fail" and the run fails. A required evaluation that does not pass also blocks the run from succeeding, even if test code caught the exception. The starter test that m3 init creates is skipped, so replace it with a real assertion before you depend on CI.

Direct tests are deterministic and need no model or credentials, so they make a good default pull-request gate. Agent tests depend on the harness, model, and approval behavior, and they incur provider cost, so keep them in a separate job or trigger when that fits your pipeline.

  • Failed assertion: the run fails, with exit code 1 for a failing pytest outcome.
  • Nothing executed: a run in which every selected test is skipped or deselected fails.
  • Required evaluation not passed: the run is not successful.
  • Missing or malformed upload credential: with --upload, the command stops with status 2 before any test runs.

Run the same suite locally and in CI

m3 ci test uses the normal project Python, storage, harness, and pytest selection. The one difference is selection: it excludes tests whose nearest M3 marker sets ci=False, which is how you keep a test that needs a person at the keyboard out of automation. Ordinary m3 test still includes that test, and paths, -k, -m, and --suite combine with the CI exclusion.

Run the credential-free selection locally before you add a workflow, so the first CI run is not the first time anyone has seen it pass. Add -n to run tests in parallel worker processes; m3 setup installs pytest-xdist for that.

Note that pytest reads testpaths only when no path is given. Paths after -- replace it, so list every directory that holds M3 tests.

Run the CI selection locally
m3 setup
m3 ci test -- tests/

# parallel worker processes
m3 ci test -n auto

GitHub Actions setup

The documented workflow runs in a consumer repository with a locked uv project: pyproject.toml and uv.lock include sf-m3[pytest,storage,judge], a .python-version file exists, and tests live under tests/. The CLI needs the pytest, storage, and judge extras in the project environment even when no test uses an LLM judge, because storage supplies the SQLite execution store that records the run.

The workflow installs the CLI at the same version as the locked SDK, runs m3 setup, then runs m3 ci test against the project’s .venv. It uses read-only repository permissions and disables checkout credential persistence. The docs pin each action to a commit SHA; use the pinned values from the guide when you copy it.

To keep reports, set M3_TIMINGS=1 and upload .m3/reports/ as an artifact with if: always(), because a failing test step would otherwise skip the upload of exactly the report you need. Set timeout-minutes on the job so a hung model call or a server that never starts cannot run for GitHub’s default 360 minutes. If your server needs compiling, build it in its own step before m3 ci test so a build failure is not reported as a test failure. The GitHub Actions guide has the complete version with these additions.

.github/workflows/m3.yml (credential-free pull request workflow from the docs)
name: M3 tests

on:
  pull_request:
  push:
    branches: [main]

permissions:
  contents: read

jobs:
  m3-tests:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1
        with:
          persist-credentials: false
      - uses: astral-sh/setup-uv@c771a70e6277c0a99b617c7a806ffedaca235ff9
        with:
          python-version-file: .python-version
      - run: uv sync --locked
      - run: uv tool install "sf-m3-cli==$(uv run --locked --no-sync python -c 'from importlib.metadata import version; print(version("sf-m3"))')"
      - run: m3 setup --python .venv/bin/python
      - run: m3 ci test --python .venv/bin/python -- tests/ -q

Credentials for agent trials in CI

The pull-request workflow above runs with no secrets. When a test puts an agent in the loop, the harness needs a provider key, and an LLM judge needs its own. M3 maps credentials by name: --credential-env codex:OPENAI_API_KEY=MY_AGENT_KEY reads the CI variable MY_AGENT_KEY and sets OPENAI_API_KEY for the Codex child process. Mappings carry variable names, not values, and a judge does not borrow an agent key.

The optional upload workflow in the docs limits secrets to trusted pushes to main and manual dispatch, installs everything before the secrets exist, and exposes them only on the final test step. It reads the agent model from a repository variable, which is not a credential. You can also select a pinned native harness version with --runtime managed when a CI test must use one.

Final step of the optional upload workflow (command from the docs)
m3 ci test --upload --python .venv/bin/python \
  --harness "codex=$M3_AGENT_MODEL" \
  --credential-env codex:OPENAI_API_KEY=MY_AGENT_KEY \
  --credential-env judge:M3_JUDGE_API_KEY=MY_JUDGE_KEY \
  -- tests/ -q

Publish runs to M3

Publishing is opt-in. A run is sent to M3 only when the command that starts it includes --upload; without it nothing is sent, and the run cannot be published later. In CI, store a CI token as the secret M3_ACCESS_TOKEN. You create it on your organization’s CI tokens page in the M3 account console, with an expiry of 7, 30, or 90 days, and the console shows the secret only once. When CI, GITHUB_ACTIONS, or GITLAB_CI has a non-empty value, M3 requires M3_ACCESS_TOKEN and does not read the interactive credential store.

M3 loads the credential before pytest starts, so a missing one stops the command before any test runs. After pytest exits 0 or 1, M3 checks the run for credential values from the test environment and publishes it if none are found. Any other exit code publishes nothing. M3 removes M3_ACCESS_TOKEN from the pytest child environment.

Uploads also need a committed m3.toml with a stable project_id, generated once by running m3 init locally. Do not run m3 init in CI. If publishing fails because the server was unreachable, the message ends with retry with m3 upload RUN_ID, and the run stays in the local database for you to retry without rerunning tests.

Frequently asked questions

How do I test an MCP server in GitHub Actions?
Use a locked uv project that includes sf-m3[pytest,storage,judge], install the matching CLI with uv tool install, run m3 setup, then run m3 ci test against the project’s .venv. The M3 GitHub Actions guide has a complete workflow you can copy.
What makes an M3 CI run fail?
A failed assertion fails the run, and so does a run in which every selected test is skipped or deselected. A required evaluation that does not pass also stops the run from succeeding. With --upload, a missing or malformed credential stops the command with status 2 before any test runs.
Do CI tests need API keys?
Direct server tests need none. Agent tests need a provider key for the harness, and an LLM judge needs its own key. Map them by variable name with --credential-env so secrets stay in the CI secret store and enter only the step that needs them.
Is publishing a run to M3 required?
No. Publishing happens only when you pass --upload, and in CI it needs an M3_ACCESS_TOKEN secret. Without --upload, nothing is sent and the run stays local.
Can I exclude a test from CI?
Yes. Set ci=False on the nearest M3 marker, and m3 ci test excludes that test and reports the count on its M3 CI excluded line. Plain m3 test still runs it.
Is M3 stable enough to gate releases?
M3 is a pre-release preview. This page was checked against release v0.2.26. Pin the CLI to the version that matches your locked SDK, as the docs workflow does, so upgrades happen when you choose.

Related

  • Run M3 tests in CI
  • Run M3 in GitHub Actions
  • Publish or retry a run
  • Manage M3 access
  • Configure credentials
  • Troubleshoot CI credentials and uploads
© 2026 SineFrame M3
HomeMCP testingMCP CI gateM3 vs MCPJamDocumentationPrivacy policyTerms of use