Open-weight MoE model matching frontier coding results on SWE-Pro and Terminal Bench for long-horizon agentic workflows at low cost.
Output Speed
Intelligence Index
Context Window
Input price
Output price
Performance
Why M2.7 Wins on Real Software Engineering, Not Just Code Completion
Benchmarks
Independent Benchmarks: Strong Coding, Uneven Beyond It
Output Speed
Intelligence Index
MMLU
GPQA
HLE
LiveCodeBench
Technical Specifications
What the model supports
Comparison
MiniMax M2.7 vs MiniMax M3: Coding Specialist or 1M-Token Frontier?
Both models come from MiniMax's current lineup and share an agentic, tool-use orientation, so teams often weigh them directly. M2.7 is the open-weight coding-and-agent specialist with a 204,800-token combined context and text-only I/O. M3 is MiniMax's provider-designated frontier model, adding a 1,000,000-token context window and multimodal input. The practical decision is whether your workload needs frontier-scale context and multimodality, or whether M2.7's proven software-engineering performance and open weights better fit a cost-sensitive, self-hostable coding stack.
Choose MiniMax M2.7 when your priority is autonomous coding agents, debugging, and productivity workflows that fit within ~200K tokens, and when open weights, self-hosting, or low per-token cost matter. Choose MiniMax M3 when you need a 1M-token context for large codebases or long agent sessions, require multimodal input, or want the provider's designated frontier model. For most focused software-engineering agents, M2.7 delivers the stronger cost-to-capability ratio; M3 earns its place when scale and modality are hard requirements.
Limitations & Trade-offs
Best-Fit Workloads
Where this model earns its place
Autonomous coding agents
M2.7's SWE-Pro score of 56.22% (matching GPT-5.3-Codex) and 55.6% on VIBE-Pro for end-to-end project delivery make it a strong engine for coding agents that plan, edit, and verify across Web, Android, and iOS tasks. Function calling and native Agent Teams support multi-agent harnesses. The always-on reasoning and slower generation mean it suits deliberate, high-value tasks more than latency-sensitive autocomplete.
Production debugging and incident response
MiniMax reports M2.7 correlating monitoring metrics with deployment timelines, performing causal reasoning on traces, and pinpointing root causes — in some internal cases reducing incident recovery to under three minutes. Its 57.0% Terminal Bench 2 result reflects system-level comprehension useful for log analysis and SRE-style workflows. Pair with real tool access to databases and repositories for the strongest results.
Office document generation and editing
M2.7 produces and performs high-fidelity multi-round edits on Word, Excel, and PowerPoint deliverables, and holds the highest open-source GDPval-AA ELO (1495). This suits automated report generation, financial modeling drafts, and structured content pipelines. Because output is text-only, downstream rendering into final office files is handled by your application layer rather than the model directly.
High-volume agentic workflows on a budget
With only ~10B active parameters per token, inference cost and latency track the active count, giving M2.7 a favorable cost-to-capability ratio for teams running many agent invocations. It occupies a low-cost tier relative to frontier proprietary coding models while retaining competitive coding accuracy — well suited to cost-sensitive, high-throughput automation where a frontier model would be prohibitively expensive at scale.