
SineFrame M3 Review: A CI Gate for MCP Servers
CodingFreeSineFrame M3 is a pytest framework for MCP servers that runs your tests through real agents like Claude Code and Codex, so a tool the agent never calls fails the build instead of shipping.
What Is SineFrame M3?
SineFrame M3 is a pytest framework for MCP servers, positioned as the CI gate between an MCP tool and the coding agents that consume it. You define tests the way you would for any Python library, but the assertions are about agent behavior: did Claude Code call your tool, with what arguments, and did the result change the agent's next move? Direct tests check schemas, results, and error handling against a live server without any model, while agent tests boot a real harness and record every MCP tool call it makes.
The Failure Mode M3 Catches
The core insight is that agents answer from memory when your tool fails to tell them not to. A registry server with a perfect get_latest_version tool is useless if the agent never calls it and just writes the oldest dependency it remembers. M3 formalizes this: run an agent against your server, assert the tool call happened with the right arguments, and fail when it did not. Fixing the tool description, rerunning, and watching the call appear is the entire loop the framework optimizes.
- Catches skipped-tool failures most schema tests miss
- Asserts tool calls happened with expected arguments
- Fail message shows observed calls versus expected
- Regression tests lock in the fix
Which Agents and Harnesses It Supports
The framework separates the test from the harness. SineFrame M3 lets the same test run against Claude Code, Codex, OpenCode, Pi, or any ACP-compatible agent, and the runner compares results side by side so you see whether a tool works in Claude Code but gets ignored by Codex. Harness configuration is per-agent and per-model, pinned via the command line, which is what makes cross-agent comparisons reproducible rather than anecdotal.
- Claude Code, Codex, OpenCode, Pi, and ACP harnesses
- Same test file runs against every harness
- Side-by-side comparison of tool usage across agents
- Models pinned per harness for reproducible runs
CI Gating and Version Comparisons
M3 ships a CI command that fails the build when an assertion breaks or when no tests executed at all — a deliberate guard against a silent, empty suite. Version pinning lets you run one suite against two server versions and compare call counts, so a refactor that stops agents from using your tool is caught before release instead of in production. The evidence generated by each run lives alongside the tests, which keeps tuning honest.
- m3 ci fails on broken assertions or zero executed tests
- Run the same suite against two versions
- Tool-call evidence generated with every run
- Hosted CI runs are on the roadmap for the preview
What M3 Adds That Other MCP Testers Miss
Generic MCP inspectors validate that a server answers requests correctly; they rarely tell you whether a real agent will actually use the tool. SineFrame M3 closes that gap by executing the agent, recording calls, and failing on silence. Its second edge is the harness matrix — the same test across Claude Code, Codex, OpenCode, and Pi catches agents whose tool-calling behavior diverges. Third, the zero-tests-fail policy makes the CI gate trustworthy instead of ceremonial.
- Agent execution beats protocol-level inspection for behavior
- Multi-harness runs expose per-agent differences
- Zero-tests-fail policy keeps the gate honest
- Evidence capture makes regressions debuggable
Platform and Setup Requirements
The CLI installs via uv, pip, or a shell installer, and a doctor command checks your harness setup. Local direct tests need no model credentials, while agent tests call your configured models through each harness's own session, so expect token costs and a few minutes per agent run. The project is in pre-release preview, which means the fastest-moving parts are the CLI flags and the CI output format.
- Install via uv, pip, or shell installer
- Direct tests need no model credentials
- Agent tests use your own model tokens
- Pre-release preview with evolving CLI conventions
Alternatives to SineFrame M3
[QAPilot MCP](/tools/qapilot-mcp) targets QA workflows through an MCP interface, while [Replay QA](/tools/replay-qa) focuses on autonomous testing of AI-built web apps rather than MCP server behavior. For building rather than testing servers, [MCP Builder AI](/tools/mcp-builder-ai) covers server creation. Among external options, the official MCP Inspector validates protocol responses, and MCPJam is the closest open-source counterpart to M3's agent-level approach, though M3's distinction is executing real harnesses with a zero-tests-fail CI gate.
- QAPilot MCP — MCP-based QA tooling
- Replay QA — autonomous testing of AI web apps
- MCP Builder AI — server creation, not testing
- External: official MCP Inspector and MCPJam
Pricing & Plans
Open source under Apache-2.0 and in pre-release preview. The CLI installs with uv, pip, or a shell script; agent harnesses need your own model credentials. Managed CI and hosted runs are on the roadmap.
Open source
The core CLI, pytest plugin, and SDK are free and open source under Apache-2.0, currently in pre-release preview.
- pytest framework for MCP servers
- Direct schema, result, and error tests without a model
- Agent tests through Claude Code, Codex, OpenCode, and Pi
- m3 ci gate that fails when zero tests execute
- Version pinning and side-by-side comparison runs
Best For
Recommended use cases and scenarios where SineFrame M3 shines.
Pros and Cons
The project earns its place in a testing stack because it tests the behavior that actually matters — whether a coding agent uses your tool — rather than whether a server responds to a probe. The honest costs are the token spend and runtime of agent tests, the pre-release maturity, and setup friction across multiple harnesses. For MCP maintainers who have shipped a tool that agents quietly ignore, M3 turns that silent failure into a red build, which is a genuinely useful inversion.
Pros
- Tests MCP servers through real agents, not just a bare model API
- Fails the build when the agent skips your tool entirely
- Plain pytest asserts with an m3 runner and evidence capture
- Runs the same tests across Claude Code, Codex, OpenCode, and Pi
- Simplifies to schema-level tests without any model credentials
Cons
- Pre-release preview, so APIs and output formats can change
- Agent tests consume model tokens and take minutes per run
- Harness setup varies by agent and provider
- Docs are early and some fields are still stabilizing
Frequently Asked Questions
Common questions about SineFrame M3, answered.
What is SineFrame M3?
SineFrame M3 is a pytest framework for MCP servers. It tests your server against real coding agents such as Claude Code, Codex, OpenCode, and Pi, recording every MCP tool call and failing the build when the agent skips your tool.
How is M3 different from the official MCP Inspector?
The MCP Inspector verifies that a server responds correctly to protocol requests. M3 goes further by executing real agent harnesses and asserting that the agent actually calls your tool — catching behavior a protocol-level test cannot see.
Does M3 require model credentials?
Direct tests for schemas, results, and errors run without any model. Agent tests boot a real harness and use your own model credentials, so they consume your tokens and take longer to run.
Is SineFrame M3 free?
Yes. It is open source under Apache-2.0 and in pre-release preview with no paid tier. Your only external cost is model API usage from agent test runs.
Which agents can M3 test against?
Claude Code, Codex, OpenCode, Pi, and any Agent Client Protocol (ACP) compatible agent. The same test file can run across all of them for side-by-side comparison.
Why would an agent skip my MCP tool?
Agents commonly answer from memory when a tool description does not clearly instruct them to check. M3 surfaces this by asserting the tool call actually happened with the expected arguments.
Reviews & Ratings
Loading reviews...
Similar Tools
More Coding tools you might like
Claude Code
Anthropic's agentic coding tool that lives in your terminal — plan, build, test, and ship software by describing tasks in plain English.
GitHub Copilot
AI coding assistant that suggests code completions and entire functions in VS Code, JetBrains, and Neovim.