MiniMax
•
MiniMax-M2
•
Released 
October 2025

MiniMax
MiniMax-M2

Open-weight 230B/10B-active MoE tuned for coding agents and long-horizon tool use at high throughput and low cost.

Modality:
Text
PDF
model ID
minimax/minimax-m2-her

Output Speed *

N/A
tok/s

Intelligence Index *

18.6
/ 100

Context Window *

65536
tokens

Input price

1.8
Anytoken

Output price

7.2
Anytoken
MiniMax-M2: Near-Frontier Coding and Agentic Execution With a 10B-Active Footprint MiniMax-M2 is an open-weight Mixture-of-Experts model from MiniMax built for coding and agentic workflows. It carries 230B total parameters but activates only about 10B per token, giving it the intelligence density of a large model with the latency and serving cost of a mid-sized one. Within MiniMax's M-series it launched as the efficient agentic release preceding M2.1, M2.5, M2.7 and M3. It fits developer agents best: multi-file edits, compile-run-fix loops, shell and browser tool chains, and long-horizon automation where fast, cheap iteration matters more than absolute peak reasoning. Access MiniMax-M2 via the AnyAPI.ai API and route your coding agents in minutes.

Performance

Why 10B Active Parameters Change the Agent Loop

MiniMax-M2 is engineered for end-to-end developer workflows: multi-file edits, coding-run-fix loops, and test-validated repairs across terminals, IDEs and CI. It generates output at 97.3 tokens per second based on MiniMax's API, well above average for open-weight models of similar size, using 230 billion total parameters with only 10 billion active during inference. That sparse activation directly shortens the plan-act-verify cycle agents depend on, so compile-test and browse-retrieve chains iterate faster at lower compute overhead. In production this favors interactive coding agents and high-volume automation where responsiveness and cost per completed task matter more than peak single-shot reasoning.

Benchmarks

MiniMax-M2 on Agentic and Coding Benchmarks

Independent testing positions MiniMax-M2 as a strong open-weight agentic model rather than an outright frontier leader. <cite index="5-10">It scores 29 on the Artificial Analysis Intelligence Index, above the median for open-weight models of similar size.</cite> <cite index="25-1,25-2">On Humanity's Last Exam it scored 31.8% with tool use, higher than many models that don't expose reasoning, outperforming Claude 4 and trailing GPT-5 only slightly.</cite> Coding-index comparisons place it below several closed frontier and larger open models but ahead of peers like GLM-4.6 and DeepSeek R1 on averaged coding evaluations. Treat these as capability signals: M2's edge is throughput-per-dollar on real agent tasks, not benchmark supremacy.

Output Speed

*
N/A
tok/s

Intelligence Index

*
18.6
/ 100

MMLU *

Broad world knowledge and problem-solving
82
%

GPQA *

PhD-level scientific reasoning across physics, biology, chemistry.
78
%

HLE *

Adherence to multi-step structured instructions.
14
%

LiveCodeBench *

Tool-calling reliability in long agentic loops.
83
%

Technical Specifications

What the model supports

MiniMax-M2 is a text in, text out MoE model. <cite index="11-6,14-1">MiniMax official documentation lists a 204,800-token context window for MiniMax-M2.</cite> This is the combined budget for system instructions, history, tool definitions, tool results, input and the model's reasoning plus answer, so large completions consume the same window as input. Critically, <cite index="20-2,20-3,20-5,20-6">M2 is an interleaved thinking model, so you must retain the assistant's <think>...</think> content in historical messages and pass it back in original format, or performance degrades.</cite> Its OpenAI- and Anthropic-compatible endpoints make integration straightforward, but the thinking-preservation requirement is the key production constraint.
Verified Specifications — 
MiniMax-M2
*
Input modalities
Text
PDF
output modalities
Text
Context window
65536
 tokens
Maximum output tokens
2048
Reasoning
No
Knowledge cutoff
October 2025
Pricing (standard)
1.8
 AnyTokens in
 / 
7.2
 AnyTokens out

Comparison

MiniMax-M2 vs GLM-4.6: Which Open-Weight Coding Agent?

MiniMax-M2 and GLM-4.6 are the two most realistic open-weight choices for coding and agent workloads, and they share near-identical context windows around 204.8K tokens. Both handle tool use and agentic execution well. <cite index="27-1,27-2">On independent testing MiniMax-M2 is faster, generating around 84.6 tokens per second versus GLM-4.6 (Reasoning) at 45.8.</cite> The practical decision is throughput and cost efficiency versus GLM-4.6's more comprehensive planning and documentation output. <cite index="32-1">GLM-4.6 is roughly 1.5x more expensive on input and 1.9x more expensive on output than MiniMax-M2.</cite>

Dimension
MiniMax-M2
GLM-4.6
Context window *
65536
tokens
204800
tokens
Output speed *
N/A
tok/s
N/A
tok/s
Intelligence Index *
18.6
14.9
Input pricing
1.8
AnyToken
2.34
AnyToken
Output pricing
7.2
AnyToken
11.4
AnyToken
Knowledge cutoff *
October 2025
September 2025

Choose MiniMax-M2 when you run long agent sessions, high-volume automation, or budget-sensitive coding loops and want faster iteration at lower cost per task. <cite index="29-7,29-8">In head-to-head coding tests the M-series delivered the same functional result as GLM at roughly half the cost, though GLM produced more comprehensive planning and documentation.</cite> Choose GLM-4.6 when you value thorough documentation, more deliberate planning, and slightly higher reliability on complex builds, and can absorb its higher price and slower generation. For pure interactive speed and cost, M2 is the stronger default.

Limitations & Trade-offs

Where MiniMax-M2 falls short

1
Legacy status within the M-series. MiniMax-M2 was the efficient October 2025 release, but MiniMax has since shipped M2.1, M2.5, M2.7 and M3, and its own catalog now labels M2 a legacy model. For greenfield projects the newer releases post materially higher agentic-coding scores, so M2 is best treated as a stable, well-understood baseline for existing workflows rather than the current default. Teams starting fresh should benchmark against a current M-series model before committing.
2
Interleaved-thinking handling is mandatory. M2 emits reasoning inside <think>...</think> blocks that must be retained and replayed in conversation history in their original format; removing them degrades quality. This adds real integration complexity for tool-use turns and increases token accounting, and naive chat wrappers that strip reasoning will silently underperform. Frameworks that don't preserve assistant thinking are a poor fit until adapted.
3
Text-only modality. MiniMax-M2 accepts and produces text only — no image, audio, or video input. Workloads needing document-image understanding, screenshots for UI agents, or multimodal retrieval must route to a different model. If your agent pipeline depends on visual inputs, M2 cannot serve that stage, and the later M3 release (with image and video input) or a multimodal alternative is required.
4
Verbose outputs can raise task-level cost. M2's low per-token price does not always translate into low cost per completed task, because agentic runs and interleaved reasoning generate high token volume. Independent notes flag high token usage as especially relevant for long-horizon agent workflows. Budget on cost-per-completed-task, not headline token rates, particularly for extended autonomous sessions.

Best-Fit Workloads

Where this model earns its place

01

Autonomous coding agents

‍
M2 is engineered for multi-file edits, compile-run-fix loops, and test-validated repairs across terminals, IDEs and CI. Its 10B active footprint keeps the plan-act-verify loop fast and cheap, so agents can iterate through debugging cycles without token cost ballooning. In practice it self-tests generated code, runs the CLI, and recovers from execution errors — well suited to CI-embedded assistants and IDE agents that must run for long autonomous stretches.

02

Long-horizon tool-use and browser agents

‍
The model plans and executes complex toolchains across shell, browser, retrieval and code runners, maintaining traceable evidence and recovering gracefully from flaky steps in BrowseComp-style evaluations. Interleaved thinking with persistent traces supports multi-step decisions across many turns. This fits research agents, web-navigation automation, and orchestration pipelines — provided the harness preserves  content between calls.

03

High-throughput, cost-sensitive automation

‍
With above-average output speed and low compute per token, M2 suits batched sampling and high-volume generation where cost per task dominates model selection. Its open-weight MIT-modified license also allows self-hosting for teams that want to control serving. Ideal for large-scale code transformation, bulk repository operations, and pipeline automation where a frontier proprietary model would be economically impractical.

04

Multilingual software engineering

‍
M2 shows strong performance on Multi-SWE-bench and SWE-bench Multilingual-style tasks, demonstrating practical effectiveness across programming languages in terminals and CI. This makes it a reasonable fit for polyglot codebases and internationalized engineering teams needing consistent behavior beyond Python-centric benchmarks — though for the highest multilingual coding accuracy, newer M-series releases score higher.

Pricing in anytokens via AnyAPI
Input
1.8
₳
Output
7.2
₳
Cache write
—
₳
Cache read
—
₳

Integration

Access MiniMax-M2 via AnyAPI.ai

Access MiniMax-M2 through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate MiniMax-M2 without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access MiniMax-M2 and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test MiniMax-M2 against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use MiniMax-M2 from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use MiniMax-M2 for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

MiniMax-M2 is an open-weight MoE model built for coding and agentic workflows — multi-file edits, compile-run-fix loops, shell and browser tool chains, and long-horizon automation. Its 230B-total/10B-active design gives near-frontier coding capability with low latency and cost, making it a strong fit for high-throughput autonomous coding agents rather than general chatbots.

MiniMax official documentation lists MiniMax-M2 with a 204,800-token context window. This is the combined budget for input and output — system instructions, conversation history, tool definitions, tool results, user input, and the model's reasoning plus answer must all fit within it, so large completions reduce available input space.

MiniMax-M2 is an interleaved thinking model that wraps reasoning in <think>...</think> blocks. You must retain and pass back this thinking content in historical messages in its original format across turns. Removing it negatively affects model performance, so agent harnesses need to preserve assistant reasoning rather than stripping it.

Yes. MiniMax-M2 is released under a Modified MIT license that allows commercial use, and weights are publicly available. This lets teams self-host or access it through hosted APIs. Note that MiniMax now labels M2 as a legacy model, superseded by M2.1, M2.5, M2.7 and M3 in its M-series.

Both are open-weight coding-agent models with roughly 204.8K-token context windows. Independent testing shows MiniMax-M2 generates tokens faster and costs less per token than GLM-4.6, making it better for high-volume, budget-sensitive agent loops. GLM-4.6 tends to produce more comprehensive planning and documentation, so it suits teams prioritizing thoroughness over speed and cost.

* Benchmark data source: Artificial Analysis artificialanalysis.ai