Qwen
•
Qwen3 Coder 480B A35B Instruct
•
Released 
July 2025

Qwen
Qwen3 Coder 480B A35B Instruct

Alibaba's largest open-weight coding MoE, built for repository-scale agentic software engineering with function calling and long context.

Modality:
Text
model ID
qwen/qwen3-coder-480b-a35b-instruct

Output Speed *

N/A
tok/s

Intelligence Index *

11.9
/ 100

Context Window *

262144
tokens

Input price

9
Anytoken

Output price

45
Anytoken
Qwen3 Coder 480B A35B: Open-Weight Agentic Coding at Repository Scale Qwen3 Coder 480B A35B is Qwen's largest coding-specialized model and the flagship of the Qwen3-Coder family from Alibaba. It is a Mixture-of-Experts model with 480B total parameters and 35B activated per token, tuned specifically for agentic software engineering: multi-turn tool use, function calling, and reasoning across entire repositories. It natively handles 256K-token context, extendable toward 1M with YaRN. Teams building autonomous coding agents, CLI-based development assistants, and large-codebase analysis benefit most, since the model was trained on long-horizon Agent RL and reaches results comparable to Claude Sonnet 4 on agentic coding tasks. Integrate Qwen3 Coder 480B A35B via the AnyAPI.ai API

Performance

Why Qwen3 Coder 480B Leads on Agentic Software Engineering

Qwen3 Coder 480B A35B is built for autonomous, multi-turn coding rather than single-shot generation. Qwen reports 67.0% on SWE-bench Verified and 61.8% on Aider Polyglot, placing it among the strongest open models and comparable to Claude Sonnet 4 on agentic coding. This matters because SWE-bench Verified rewards planning, tool invocation, and iterative debugging across real repositories, not isolated snippets. In production, that translates to a model capable of driving CLI agents, cross-file refactors, and pull-request generation. Its sparse MoE design activates only 35B of 480B parameters, keeping inference throughput usable at flagship-level quality.

Benchmarks

Qwen3 Coder 480B Benchmark Interpretation

On the Artificial Analysis Intelligence Index the model scores 18, above the median for comparable open-weight non-reasoning models. Independent measurement puts output at roughly 68 tokens per second and time to first token around 3.0s on Alibaba's API — throughput above average, but a relatively high TTFT for latency-sensitive UIs. On coding-specific benchmarks it is stronger: 67.0% SWE-bench Verified and 61.8% Aider Polyglot as reported by Qwen. The takeaway: this is a coding and agentic specialist rather than a general-intelligence leader, and its scores are most meaningful when interpreted against software-engineering workloads.

Output Speed

*
N/A
tok/s

Intelligence Index

*
11.9
/ 100

MMLU *

Broad world knowledge and problem-solving
79
%

GPQA *

PhD-level scientific reasoning across physics, biology, chemistry.
62
%

HLE *

Adherence to multi-step structured instructions.
5
%

LiveCodeBench *

Tool-calling reliability in long agentic loops.
59
%

Technical Specifications

What the model supports

Qwen3 Coder 480B A35B is a text-and-code model — it accepts text, code, and structured function calls, and returns text and code only; there is no image, audio, or video modality. It exposes native function calling and tool choice with a purpose-built call format, making it suitable for agent scaffolds. Critically, it runs in non-thinking mode only and does not emit reasoning blocks, so there are no extended-thinking controls. The 256K native context is the standout production feature, enabling whole-repository prompts without chunking, though hosted endpoints may cap maximum output well below the model default.
Verified Specifications — 
Qwen3 Coder 480B A35B Instruct
*
Input modalities
Text
output modalities
Text
Context window
262144
 tokens
Maximum output tokens
0
Reasoning
No
Knowledge cutoff
July 2025
Pricing (standard)
9
 AnyTokens in
 / 
45
 AnyTokens out

Comparison

Qwen3 Coder 480B A35B vs Kimi K2 Instruct for Coding Agents

Both are large open-weight MoE models positioned for agentic coding, and both land near Claude Sonnet 4 territory on SWE-bench Verified. Kimi K2 Instruct (0905) scores about 69.2% on SWE-bench Verified versus Qwen3 Coder's 67.0%, with the two trading places across tool-use benchmarks like τ²-bench. The realistic decision is not which model writes better snippets — both are strong — but which integrates more cleanly into your agent scaffold, tool-calling format, and context needs. Qwen3 Coder's native 256K context, extendable toward 1M, and its purpose-built function-call format are the deciding factors for repository-scale workflows.

Dimension
Qwen3 Coder 480B A35B Instruct
Kimi K2 Thinking
Context window *
262144
tokens
262144
tokens
Output speed *
N/A
tok/s
0.00
tok/s
Intelligence Index *
11.9
22
Input pricing
9
AnyToken
3.42
AnyToken
Output pricing
45
AnyToken
13.8
AnyToken
Knowledge cutoff *
July 2025
November 2025

Choose Qwen3 Coder 480B A35B when you need repository-scale context (256K native, up to ~1M with YaRN), a coding-first model with a well-supported agentic call format across tools like Qwen Code and Cline, and open weights you can self-host. Choose Kimi K2 Instruct when you want marginally higher raw SWE-bench Verified numbers, prefer its tool-use profile, or already run its ecosystem. For most teams the choice hinges on context length and scaffold compatibility rather than headline benchmark deltas.

Limitations & Trade-offs

Where Qwen3 Coder 480B A35B Instruct falls short

1
No reasoning mode. Qwen3 Coder 480B A35B runs in non-thinking mode only and does not generate extended-thinking blocks. For tasks that benefit from visible step-by-step deliberation — hard algorithmic reasoning, math-heavy planning — a dedicated reasoning model may perform better. In agent scaffolds you must supply explicit planning prompts rather than relying on internal chain-of-thought, which shifts more orchestration responsibility onto your framework.
2
High time to first token. Independent measurement reports TTFT around 3.0 seconds on Alibaba's API, at the higher end for its class. Combined with a large-model warm-up profile, this makes the model less suited to latency-sensitive interactive UIs like inline autocomplete. It is better matched to asynchronous agent runs and batch code tasks where total throughput matters more than first-token responsiveness.
3
Text and code only. There is no image, audio, or video input, so it cannot process screenshots, design mockups, diagrams, or UI states directly. Multimodal agentic workflows — debugging from a screenshot or reading a rendered chart — require a separate vision-capable model. Verify modality assumptions before building it into pipelines that expect visual context.
4
Cost profile for a non-reasoning model. Independent analysis flags the model as comparatively expensive among open-weight non-reasoning models of similar size, and hosted pricing rises for requests above the long-context threshold. For very high-volume, simple code generation, a smaller sibling such as Qwen3-Coder-30B-A3B can be more economical while retaining agentic capability.

Best-Fit Workloads

Where this model earns its place

01

Autonomous coding agents

‍
The model was post-trained with long-horizon Agent RL and reaches ~67% on SWE-bench Verified, which rewards planning, tool use, and iterative debugging. It integrates with agent scaffolds like Qwen Code and Cline via a purpose-built function-call format, making it a strong engine for autonomous engineering agents that reproduce bugs, edit multiple files, run tests, and open pull requests. Pair it with a robust orchestration layer since it has no internal reasoning mode.

02

Repository-scale code understanding

‍
With a native 256K-token context extendable toward 1M via YaRN, the model can ingest large portions of a codebase in a single prompt. This suits whole-repository analysis, cross-file refactoring, dependency tracing, and generating architecture documentation without heavy chunking. Long context reduces retrieval complexity for RAG-style code assistants, though throughput and cost rise at the top of the context range.

03

Multi-language code generation and migration

‍
Trained heavily on source code across many languages, the model scored 61.8% on Aider Polyglot, a multi-language edit benchmark. That makes it a practical choice for polyglot codebases, framework migrations, and translating logic between languages such as Python, JavaScript, Java, C++, Go, and Rust. Its diff-editing competence supports patch-style workflows rather than only full-file rewrites.

04

Tool-use and CLI automation

‍
Native function calling and tool choice, combined with agentic RL training on browser and CLI environments, let the model drive terminal-based automation and external-tool pipelines. It performs competitively on tool-use benchmarks, making it suitable for build/test automation agents and environment-interacting workflows. Because TTFT is relatively high, prefer asynchronous execution over interactive real-time loops.

Pricing in anytokens via AnyAPI
Input
9
₳
Output
45
₳
Cache write
—
₳
Cache read
—
₳

Integration

Access Qwen3 Coder 480B A35B Instruct via AnyAPI.ai

Access Qwen3 Coder 480B A35B Instruct through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate Qwen3 Coder 480B A35B Instruct without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access Qwen3 Coder 480B A35B Instruct and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test Qwen3 Coder 480B A35B Instruct against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use Qwen3 Coder 480B A35B Instruct from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use Qwen3 Coder 480B A35B Instruct for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

Qwen3 Coder 480B A35B supports 262,144 tokens (256K) natively, extendable toward roughly 1 million tokens using YaRN extrapolation. This lets it process large portions of a codebase in one prompt. Note that some hosted endpoints cap maximum output separately — commonly around 32,768 to 65,536 tokens — which is distinct from the input context window.

No. The model runs in non-thinking mode only and does not generate extended-thinking or reasoning blocks in its output. Its agentic strength comes from long-horizon reinforcement learning on tool use and multi-turn coding, not visible chain-of-thought. For tasks needing explicit deliberation, supply planning steps in your prompt or use a dedicated reasoning model.

Qwen reports about 67.0% on SWE-bench Verified and 61.8% on Aider Polyglot, placing it among the strongest open-weight coding models and comparable to Claude Sonnet 4 on agentic coding tasks. On the Artificial Analysis Intelligence Index it scores 18, above the median for its class. Its results are most meaningful for software-engineering workloads rather than general intelligence.

Yes. It is an open-weight model released by Qwen (Alibaba) and can be self-hosted or accessed through multiple providers with OpenAI-compatible endpoints, including Alibaba Cloud Model Studio, AWS Bedrock, and Google Vertex AI. Alibaba also open-sourced Qwen Code, a CLI tool designed to drive the model's agentic coding capabilities.

No. It accepts only text, code, and structured function calls, and returns text and code. There is no image, audio, or video input. Multimodal workflows such as debugging from a screenshot require a separate vision-capable model alongside it.

* Benchmark data source: Artificial Analysis artificialanalysis.ai