Near-flagship agentic coding and computer-use performance at Sonnet-tier pricing, with a 1M-token context window.
Output Speed *
Intelligence Index *
Context Window *
Input price
Output price
Performance
Where Sonnet 4.6 Matches the Flagship Tier
Benchmarks
Independent Benchmarks: Coding, Computer Use, and Reasoning
Output Speed
Intelligence Index
MMLU *
GPQA *
HLE *
LiveCodeBench *
Technical Specifications
What the model supports
Limitations & Trade-offs
Best-Fit Workloads
Where this model earns its place
Agentic coding assistants
Sonnet 4.6's 79.6% on SWE-bench Verified and mature tool calling—including code execution, tool search, and memory—make it well suited to coding agents that read repositories, plan multi-file changes, run tests, and iterate on failures. Developers reportedly preferred it over Sonnet 4.5 in 70% of Claude Code comparisons, citing better instruction following and fewer hallucinations. The 1M context (beta) supports full-codebase analysis in a single request. Validate output latency for interactive IDE integrations.
Computer-use and browser automation
At 72.5% on OSWorld-Verified, Sonnet 4.6 performs near Opus 4.6 on autonomous desktop and browser tasks—navigating GUIs, filling multi-step forms, and coordinating across tabs. This suits workflow automation over legacy software without APIs and QA agents that operate through the screen. Anthropic also reports improved resistance to prompt injection versus Sonnet 4.5, which matters when agents act autonomously on untrusted content.
Long-context document and codebase analysis
The 1M-token context window (beta) lets Sonnet 4.6 ingest entire codebases, long contract sets, or large document collections in one prompt, reducing chunking complexity in RAG pipelines. Combined with beta context compaction that summarizes older context, this supports extended conversations and multi-document reasoning. Because the 1M window is beta, confirm availability and pricing terms before designing a pipeline that depends on it.
Office and finance-agent automation
Sonnet 4.6 leads real-world office productivity tasks (reported 1633 Elo on GDPval-AA) and scores strongly on the Finance Agent benchmark (63.3%), in some cases edging Opus 4.6. This makes it a fit for agents that produce polished documents, perform financial modeling, and run compliance-style review. For enterprise deployments processing high volumes of these tasks, near-flagship accuracy at Sonnet pricing is the core value proposition.