Anthropic
•
0
•
Released 
February 2026

Anthropic
0

Near-flagship agentic coding and computer-use performance at Sonnet-tier pricing, with a 1M-token context window.

Modality:
Text
Image
PDF
model ID
anthropic/claude-sonnet-4.6

Output Speed *

0.00
tok/s

Intelligence Index *

0
/ 100

Context Window *

0
tokens

Input price

18
Anytoken

Output price

90
Anytoken
Claude Sonnet 4.6: Flagship-Class Agentic Coding at Sonnet Pricing Claude Sonnet 4.6 is Anthropic's mid-tier model in the Claude 4 family, positioned between Haiku and Opus. It closes most of the gap to Opus-class models on the workloads that dominate production: agentic coding, computer use, and real-world office tasks. On coding and computer-use benchmarks it lands within a point or two of Opus 4.6 while staying in the Sonnet cost tier. Add a 1M-token context window (beta), configurable thinking effort, and mature tool calling, and Sonnet 4.6 becomes the default choice for coding agents and long-horizon automation that previously required a premium model. Start building with Claude Sonnet 4.6 via the AnyAPI.ai API.

Performance

Where Sonnet 4.6 Matches the Flagship Tier

Sonnet 4.6's strongest area is agentic coding and computer use. On SWE-bench Verified it reaches 79.6% under announcement conditions, and on OSWorld-Verified it scores 72.5% for autonomous desktop navigation. Both land within roughly one to two points of Opus 4.6, yet Sonnet 4.6 sits in the lower Sonnet cost tier. For teams running coding assistants, browser automation, or multi-step office workflows, this means near-flagship task completion without flagship economics. The practical consequence: most agentic and coding workloads no longer justify reaching for a premium Opus-class model, letting teams route high-volume agent traffic through Sonnet 4.6.

Benchmarks

Independent Benchmarks: Coding, Computer Use, and Reasoning

On Artificial Analysis, Claude Sonnet 4.6 with adaptive reasoning at max effort scores 48 on the Intelligence Index, well above the reasoning-model median of 34. Independent and Anthropic-reported figures put it at 79.6% on SWE-bench Verified and 72.5% on OSWorld-Verified, close to Opus 4.6 on both. Its largest single-generation jump is on ARC-AGI-2, moving to roughly 58% from Sonnet 4.5's 13.6%. Note that Artificial Analysis characterizes the reasoning variant as comparatively slow on output speed, so latency-sensitive workloads should validate throughput before committing.

Output Speed

*
0.00
tok/s

Intelligence Index

*
0
/ 100

MMLU *

Broad world knowledge and problem-solving
0
%

GPQA *

PhD-level scientific reasoning across physics, biology, chemistry.
0
%

HLE *

Adherence to multi-step structured instructions.
0
%

LiveCodeBench *

Tool-calling reliability in long agentic loops.
0
%

Technical Specifications

What the model supports

Sonnet 4.6 accepts text and image input and returns text output; it is not an audio or image-generation model. It offers a 1M-token context window in beta alongside the standard tier, with a maximum output of 128,000 tokens. The 1M context is its most consequential specification for developers—it allows full-codebase analysis and long multi-document workflows in a single request, previously an Opus-only capability. Configurable thinking effort (adaptive and extended thinking) lets teams trade latency for reasoning depth per request, which matters when balancing agent responsiveness against harder planning tasks.
Verified Specifications — 
0
*
Input modalities
Text
Image
PDF
output modalities
Text
Context window
0
 tokens
Maximum output tokens
0
Reasoning
No
Knowledge cutoff
February 2026
Pricing (standard)
18
 AnyTokens in
 / 
90
 AnyTokens out

Limitations & Trade-offs

Where 0 falls short

1
Output speed and latency. Artificial Analysis characterizes the Sonnet 4.6 reasoning variant as notably slow on output tokens per second relative to comparable models. Extended thinking adds further latency because reasoning tokens are generated before the answer. For interactive, latency-sensitive applications—live chat, autocomplete, real-time UX—this can hurt responsiveness. Teams needing fast turnaround should lower thinking effort, disable extended thinking, or consider Claude Haiku 4.5 for speed-critical paths.
2
Long-output economics. Output pricing sits in the premium Sonnet tier and is materially higher than input pricing. Workloads that generate very long responses—large code files, extensive documents, verbose agent transcripts—accumulate cost quickly, and the 128,000-token max output makes long generations possible. For high-volume, long-form generation where accuracy tolerances are looser, a cheaper model or tighter output constraints may be more economical.
3
1M context is beta, not the default tier. The 1M-token context window is offered in beta and accompanied by beta context compaction; the standard tier is 200,000 tokens. Anthropic has previously retired 1M-token betas for earlier Sonnet models. Teams building long-context RAG or full-codebase workflows should treat the 1M window as a capability that may carry different availability or pricing terms rather than a guaranteed permanent default.
4
Deep-reasoning ceiling below Opus. Despite a large ARC-AGI-2 jump, Sonnet 4.6 still trails Opus 4.6 on the hardest reasoning benchmarks—roughly 58% vs 69% on ARC-AGI-2 and a significant gap on GPQA Diamond graduate-level science. For workloads dominated by novel abstraction or expert scientific reasoning rather than coding and agents, the flagship remains the stronger choice.

Best-Fit Workloads

Where this model earns its place

01

Agentic coding assistants

‍
Sonnet 4.6's 79.6% on SWE-bench Verified and mature tool calling—including code execution, tool search, and memory—make it well suited to coding agents that read repositories, plan multi-file changes, run tests, and iterate on failures. Developers reportedly preferred it over Sonnet 4.5 in 70% of Claude Code comparisons, citing better instruction following and fewer hallucinations. The 1M context (beta) supports full-codebase analysis in a single request. Validate output latency for interactive IDE integrations.

02

Computer-use and browser automation

‍
At 72.5% on OSWorld-Verified, Sonnet 4.6 performs near Opus 4.6 on autonomous desktop and browser tasks—navigating GUIs, filling multi-step forms, and coordinating across tabs. This suits workflow automation over legacy software without APIs and QA agents that operate through the screen. Anthropic also reports improved resistance to prompt injection versus Sonnet 4.5, which matters when agents act autonomously on untrusted content.

03

Long-context document and codebase analysis

‍
The 1M-token context window (beta) lets Sonnet 4.6 ingest entire codebases, long contract sets, or large document collections in one prompt, reducing chunking complexity in RAG pipelines. Combined with beta context compaction that summarizes older context, this supports extended conversations and multi-document reasoning. Because the 1M window is beta, confirm availability and pricing terms before designing a pipeline that depends on it.

04

Office and finance-agent automation

‍
Sonnet 4.6 leads real-world office productivity tasks (reported 1633 Elo on GDPval-AA) and scores strongly on the Finance Agent benchmark (63.3%), in some cases edging Opus 4.6. This makes it a fit for agents that produce polished documents, perform financial modeling, and run compliance-style review. For enterprise deployments processing high volumes of these tasks, near-flagship accuracy at Sonnet pricing is the core value proposition.

Pricing in anytokens via AnyAPI
Input
18
₳
Output
90
₳
Cache write
22.5
₳
Cache read
1.8
₳

Integration

Access 0 via AnyAPI.ai

Access 0 through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate 0 without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access 0 and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test 0 against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use 0 from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use 0 for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

Claude Sonnet 4.6 is strongest at agentic coding and computer use. It scores 79.6% on SWE-bench Verified and 72.5% on OSWorld-Verified—within one to two points of Anthropic's Opus 4.6 flagship—while staying in the lower Sonnet cost tier. It also leads several real-world office and finance-agent tasks, making it a default for coding assistants and workflow automation.

Claude Sonnet 4.6 offers a 1M-token context window in beta, alongside a standard 200,000-token tier. Maximum output is 128,000 tokens. The 1M window enables full-codebase and multi-document analysis in a single request, but because it is beta, developers should confirm current availability and terms before depending on it in production.

Yes. Claude Sonnet 4.6 supports adaptive and extended thinking with configurable effort, and Anthropic reports strong performance even with extended thinking off. It offers mature tool calling, including programmatic tool calling, tool search, code execution, memory, structured outputs, and web search/fetch tools, making it well suited to autonomous agent workflows.

On coding and computer use, Sonnet 4.6 is within about a point of Opus 4.6 (79.6% vs 80.8% on SWE-bench Verified; 72.5% vs 72.7% on OSWorld) at roughly one-fifth the cost. Opus 4.6 leads clearly on deep reasoning benchmarks such as ARC-AGI-2 and GPQA Diamond. Choose Sonnet 4.6 for cost-effective agents and coding; Opus 4.6 for the hardest reasoning.

Yes. Claude Sonnet 4.6 accepts both text and image input and produces text output. It can analyze and answer questions about images, which supports computer-use and document workflows. It does not generate images or process audio—it is a text-output model with vision input, not a full multimodal generation model.

* Benchmark data source: Artificial Analysis artificialanalysis.ai