Qwen
•
Qwen3 Coder Next
•
Released 
February 2026

Qwen
Qwen3 Coder Next

An 80B/3B-active MoE coding model that delivers frontier-level agentic performance at a fraction of the inference cost.

Modality:
Text
PDF
model ID
qwen/qwen3-coder-next

Output Speed *

109.01
tok/s

Intelligence Index *

9.2
/ 100

Context Window *

262144
tokens

Input price

0.72
Anytoken

Output price

4.5
Anytoken
Qwen3 Coder Next: Frontier Agentic Coding on a 3B Active Footprint Qwen3-Coder-Next is an open-weight coding model from Alibaba's Qwen team, built on Qwen3-Next-80B-A3B with hybrid attention and a sparse Mixture-of-Experts design. It carries 80B total parameters but activates only 3B per token, so it delivers coding and agentic quality closer to models with 10–20x more active compute while staying cheap to run. Trained heavily on executable coding tasks and reinforcement learning, it targets long-horizon coding agents, CLI/IDE tools, and repository-scale work. It runs exclusively in non-thinking mode, simplifying integration into production agent scaffolds. Start building coding agents with Qwen3-Coder-Next via the AnyAPI.ai API.

Performance

Agentic Coding Quality Without the Active-Parameter Cost

Qwen3-Coder-Next is built for long-horizon coding agents: multi-turn tool use, editing across a repository, and recovering from execution failures. Its strongest evidence is SWE-Bench Verified, where Qwen reports 70.6% using the SWE-Agent scaffold, competitive with substantially larger open-weight models. Because only 3B of 80B parameters activate per token, that quality arrives with low inference cost and high throughput (roughly 143 tokens/sec median across providers). For teams running always-on coding agents where every request incurs compute, this efficiency-to-quality ratio directly reduces the cost of continuous automated software work.

Benchmarks

How Qwen3-Coder-Next Scores on Independent Agentic Benchmarks

On the Qwen technical report, Qwen3-Coder-Next reaches 70.6% on SWE-Bench Verified with the SWE-Agent scaffold, ahead of DeepSeek-V3.2 (70.2%) on the same chart and competitive on SWE-Bench Multilingual, SWE-Bench Pro, Terminal-Bench 2.0, and Aider. Independently, Artificial Analysis places it at 21 on its Intelligence Index — well above the median for comparable open-weight non-reasoning models — and measures around 143 tokens/sec output speed. Artificial Analysis also flags it as notably verbose, generating far more tokens than the median on Intelligence Index tasks.

Output Speed

*
109.01
tok/s

Intelligence Index

*
9.2
/ 100

MMLU *

Broad world knowledge and problem-solving
0
%

GPQA *

PhD-level scientific reasoning across physics, biology, chemistry.
74
%

HLE *

Adherence to multi-step structured instructions.
10
%

LiveCodeBench *

Tool-calling reliability in long agentic loops.
0
%

Technical Specifications

What the model supports

Qwen3-Coder-Next is a text-in, text-out model with a native 256K-token context window, enough to hold large repositories, full PR diffs, and long API specifications. It exposes function calling and JSON-schema structured outputs through OpenAI-compatible tooling. The most consequential detail for production is that it runs only in non-thinking mode and emits no reasoning blocks — predictable output structure that simplifies agent integration, but no adjustable reasoning depth. Note that maximum output tokens vary by host; the Amazon Bedrock card lists 16K while some OpenRouter providers advertise higher completion limits.
Verified Specifications — 
Qwen3 Coder Next
*
Input modalities
Text
PDF
output modalities
Text
Context window
262144
 tokens
Maximum output tokens
262144
Reasoning
No
Knowledge cutoff
February 2026
Pricing (standard)
0.72
 AnyTokens in
 / 
4.5
 AnyTokens out

Comparison

Qwen3-Coder-Next vs Qwen3-Coder 480B A35B: Which Coding Model?

Both models come from Qwen's coding family, target agentic software engineering, and share a 256K native context window with OpenAI-compatible tool calling and structured outputs. The difference is scale and economics. Qwen3-Coder 480B A35B activates 35B parameters per token from a 480B total, while Qwen3-Coder-Next activates only 3B from an 80B total. The practical decision is whether you need the raw ceiling of the larger flagship or the far cheaper, faster, self-hostable footprint of Next for high-volume, always-on agent workloads.

Dimension
Qwen3 Coder Next
Qwen3 Coder 480B A35B Instruct
Context window *
262144
tokens
262144
tokens
Output speed *
109.01
tok/s
N/A
tok/s
Intelligence Index *
9.2
11.9
Input pricing
0.72
AnyToken
9
AnyToken
Output pricing
4.5
AnyToken
45
AnyToken
Knowledge cutoff *
February 2026
July 2025

Choose Qwen3-Coder-Next when you run continuous coding agents, want low per-request cost, need to self-host on modest hardware, or value high throughput on routine software tasks. Choose Qwen3-Coder 480B A35B when you need maximum capability headroom on the hardest problems and can afford the larger active-parameter compute. For most everyday repository work, Next closes much of the quality gap while being dramatically cheaper to operate; reserve the 480B model for the most demanding cases.

Limitations & Trade-offs

Where Qwen3 Coder Next falls short

1
Non-thinking mode only. Qwen3-Coder-Next runs exclusively in non-thinking mode and never emits reasoning blocks, with no adjustable reasoning depth. This simplifies agent integration and keeps output predictable, but it removes extended-thinking as a lever for the hardest, multi-step reasoning problems. Workloads that benefit from explicit chain-of-thought scaling — complex algorithmic design or deep debugging — may see a ceiling here. If you need controllable reasoning effort, a dedicated reasoning model is a better fit.
2
High verbosity. Artificial Analysis flags Qwen3-Coder-Next as very verbose, generating roughly 32M tokens on Intelligence Index tasks against a ~20M median. Since output tokens drive cost and latency, verbose generation partially offsets the model's low per-token pricing on long tasks. For cost-sensitive, high-volume agent loops, teams should budget for higher output volume and consider tighter prompting or output constraints to keep spend and response times in check.
3
Text-only, no multimodality. The model accepts and produces only text and code — no image, audio, video, or document image input. Coding agents that need to interpret screenshots, UI mockups, diagrams, or scanned specs cannot use Qwen3-Coder-Next for that step and must pair it with a vision-capable model. If your workflow depends on visual input, choose a multimodal model instead.
4
Variable max output across hosts. Maximum output tokens differ by provider — the Amazon Bedrock card lists 16K while some OpenRouter providers advertise much higher completion limits. This inconsistency matters for tasks that generate very long files or large multi-file diffs, where a low output cap forces chunking. Verify the specific limit of your chosen host before designing an agent around long single-shot generations.

Best-Fit Workloads

Where this model earns its place

01
Autonomous coding agents Qwen3-Coder-Next was trained specifically for long-horizon agentic coding — multi-turn tool use, editing across files, and recovering from execution failures. Its 70.6% SWE-Bench Verified score on the SWE-Agent scaffold demonstrates real-world GitHub-issue resolution ability. Combined with cheap per-token economics and high throughput, it suits always-on agents that run continuously in CLI and IDE tools. The main caveat is verbosity, which can raise cost on very long agent loops.
02

Repository-scale understanding

‍
The native 256K-token context window is large enough to load substantial codebases, full pull-request diffs, and lengthy API specifications into a single request. This supports code review, cross-file refactoring, and questions that span many files without aggressive retrieval chunking. For repositories exceeding the window, you'll still need retrieval or file selection, but for most mid-sized projects the context comfortably covers the working set.

03

Cost-sensitive, high-volume automation

‍
With only 3B active parameters, Qwen3-Coder-Next delivers coding quality comparable to much larger models at a fraction of the inference cost, making it attractive for high-frequency automation such as CI code fixes, batch refactoring, or generating boilerplate at scale. Its fast output speed keeps throughput high. Because it is open-weight under a permissive license, teams can also self-host to eliminate per-token API spend for large, sustained workloads.

04

Local and self-hosted development

‍
As an open-weight model, Qwen3-Coder-Next can run locally on high-end consumer and single-GPU setups, fitting privacy-sensitive and offline development where code cannot leave the network. It adapts to common scaffolds like Claude Code, Qwen Code, and Cline. This makes it a fit for regulated environments and IP-sensitive teams, though local throughput depends heavily on available VRAM and quantization.

Pricing in anytokens via AnyAPI
Input
0.72
₳
Output
4.5
₳
Cache write
—
₳
Cache read
—
₳

Integration

Access Qwen3 Coder Next via AnyAPI.ai

Access Qwen3 Coder Next through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate Qwen3 Coder Next without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access Qwen3 Coder Next and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test Qwen3 Coder Next against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use Qwen3 Coder Next from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use Qwen3 Coder Next for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

Qwen3-Coder-Next is a coding model built for agentic software engineering: multi-turn tool use, editing across a repository, and recovering from execution failures. It reports 70.6% on SWE-Bench Verified with the SWE-Agent scaffold. Its 3B active parameters make it especially suited to cost-sensitive, high-volume coding agents in CLI and IDE tools.

Qwen3-Coder-Next has a native 256K-token context window (262,144 tokens). That is large enough to load substantial codebases, full pull-request diffs, and long API specifications into a single request, supporting repository-scale understanding and cross-file refactoring without aggressive retrieval chunking.

No. Qwen3-Coder-Next runs exclusively in non-thinking mode and does not emit reasoning blocks, and there is no adjustable reasoning-effort control. This makes its output predictable and easy to integrate into agent pipelines, but it means extended-thinking is not available for the hardest multi-step problems.

Yes. Qwen3-Coder-Next is an open-weight model released by Alibaba's Qwen team on February 4, 2026, built on Qwen3-Next-80B-A3B with hybrid attention and a Mixture-of-Experts design. Being open-weight, it can be self-hosted for privacy-sensitive or offline development in addition to being accessed via API.

Independent testing by Artificial Analysis measures Qwen3-Coder-Next at roughly 143 tokens per second output speed as a median across providers, which is notably fast for its class. Individual hosts vary — some FP8 providers exceed 170 tokens per second. Note the model is also flagged as verbose, so long tasks generate many output tokens.

* Benchmark data source: Artificial Analysis artificialanalysis.ai