Anthropic
•
Claude 4 Sonnet (Non-reasoning)
•
Released 
May 2025

Anthropic
Claude 4 Sonnet (Non-reasoning)

Anthropic's balanced hybrid-reasoning model with frontier coding performance and optional extended thinking for agentic workflows.

Modality:
Text
Image
PDF
model ID
anthropic/claude-sonnet-4

Output Speed *

0.00
tok/s

Intelligence Index *

16.6
/ 100

Context Window *

200000
tokens

Input price

18
Anytoken

Output price

90
Anytoken

Claude Sonnet 4: Frontier Coding and Agentic Workflows at Balanced Cost

Claude Sonnet 4 is Anthropic's balanced-tier hybrid reasoning model, released in May 2025 as the direct successor to Sonnet 3.7. It sits below the Opus flagship but brings frontier coding and agentic performance to high-volume production use. Sonnet 4 can switch between fast standard responses and an extended thinking mode for step-by-step reasoning. Its defining strength is real-world software engineering: it scores strongly on SWE-bench Verified and handles autonomous codebase navigation, multi-step tool use, and code review. Coding assistants and long-running agents benefit most from its capability-to-cost balance.

Start building with Claude Sonnet 4 through the AnyAPI.ai unified API.

Performance

Where Claude Sonnet 4 Earns Its Place: Real-World Software Engineering

Claude Sonnet 4 is built for practical coding work rather than synthetic puzzles. Anthropic reports a 72.7% score on SWE-bench Verified, a benchmark drawn from real GitHub issues where a model must understand a repository, implement a fix, and pass existing tests. Anthropic also reports that navigation errors during codebase exploration dropped to near zero versus Sonnet 3.7. For production teams this means fewer dead-end patches and more reliable autonomous edits, which directly reduces the human review burden in coding assistants and agentic pipelines that delegate real bugfix and refactoring tasks.

Benchmarks

Claude Sonnet 4 Benchmarks: Coding Strength, Mid-Tier Intelligence Index

Independent evaluation from Artificial Analysis places Claude Sonnet 4 (non-reasoning) at 26 on its composite Intelligence Index, above the comparable median of 21 but below reasoning-heavy frontier models. On coding, Anthropic's official SWE-bench Verified result of 72.7% reflects strong real-world software engineering, with higher scores achievable using parallel test-time compute. The pattern is consistent: Sonnet 4 is a coding and agentic workhorse rather than a maximum-reasoning model. Teams needing the deepest reasoning on math or research-heavy tasks should evaluate a reasoning-optimized alternative, while coding-first workloads see Sonnet 4's advantage clearly.

Output Speed

*
0.00
tok/s

Intelligence Index

*
16.6
/ 100

MMLU *

Broad world knowledge and problem-solving
84
%

GPQA *

PhD-level scientific reasoning across physics, biology, chemistry.
68
%

HLE *

Adherence to multi-step structured instructions.
4
%

LiveCodeBench *

Tool-calling reliability in long agentic loops.
45
%

Technical Specifications

What the model supports

Claude Sonnet 4 is a proprietary hybrid reasoning model accepting text and image input and returning text. It ships with a 200,000-token context window by default, with a 1M-token context available in public beta. The two production-critical constraints are separate: the large input context does not raise the 64,000-token maximum output, so extremely long single-response generation still needs chunking. Extended thinking is optional, letting teams trade latency for deeper reasoning per request. Function calling and structured tool use are supported, making Sonnet 4 suitable for agentic pipelines that maintain state across many tool calls.
Verified Specifications — 
Claude 4 Sonnet (Non-reasoning)
*
Input modalities
Text
Image
PDF
output modalities
Text
Context window
200000
 tokens
Maximum output tokens
64000
Reasoning
Yes
Knowledge cutoff
May 2025
Pricing (standard)
18
 AnyTokens in
 / 
90
 AnyTokens out

Quickstart

Sample code for Claude 4 Sonnet (Non-reasoning)

import requests

url = "https://api.anyapi.ai/v1/chat/completions"

payload = {
    "model": "claude-4-sonnet",
    "messages": [
        {
            "role": "user",
            "content": "Hello"
        }
    ]
}
headers = {
    "Authorization": "Bearer  AnyAPI_API_KEY",
    "Content-Type": "application/json"
}

response = requests.post(url, json=payload, headers=headers)

print(response.json())

‍

import requests url = "https://api.anyapi.ai/v1/chat/completions" payload = { "model": "claude-4-sonnet", "messages": [ { "role": "user", "content": "Hello" } ] } headers = { "Authorization": "Bearer AnyAPI_API_KEY", "Content-Type": "application/json" } response = requests.post(url, json=payload, headers=headers) print(response.json())
View docs
Copy
Code is copied
const url = 'https://api.anyapi.ai/v1/chat/completions';
const options = {
  method: 'POST',
  headers: {Authorization: 'Bearer  AnyAPI_API_KEY', 'Content-Type': 'application/json'},
  body: '{"model":"claude-4-sonnet","messages":[{"role":"user","content":"Hello"}]}'
};

try {
  const response = await fetch(url, options);
  const data = await response.json();
  console.log(data);
} catch (error) {
  console.error(error);
}
const url = 'https://api.anyapi.ai/v1/chat/completions'; const options = { method: 'POST', headers: {Authorization: 'Bearer AnyAPI_API_KEY', 'Content-Type': 'application/json'}, body: '{"stream":false,"tool_choice":"auto","model":"claude-4-sonnet","messages":[{"role":"user","content":"Test prompt"}]}' }; try { const response = await fetch(url, options); const data = await response.json(); console.log(data); } catch (error) { console.error(error); }
View docs
Copy
Code is copied
curl --request POST \
  --url https://api.anyapi.ai/v1/chat/completions \
  --header 'Authorization: Bearer  AnyAPI_API_KEY' \
  --header 'Content-Type: application/json' \
  --data '{
  "model": "claude-4-sonnet",
  "messages": [
    {
      "role": "user",
      "content": "Hello"
    }
  ]
}'
curl --request POST \ --url https://api.anyapi.ai/v1/chat/completions \ --header 'Authorization: Bearer AnyAPI_API_KEY' \ --header 'Content-Type: application/json' \ --data '{ "model": "claude-4-sonnet", "messages": [ { "role": "user", "content": "Hello" } ] }'
View docs
Copy
Code is copied
View docs
Code examples coming soon...

Comparison

Claude Sonnet 4 vs Claude Opus 4: Which Tier for Your Workload?

Claude Sonnet 4 and Claude Opus 4 launched together in the Claude 4 family in May 2025 and share the same hybrid reasoning architecture, extended thinking mode, tool use, and 200K default context window. Both lead on SWE-bench Verified, with scores in a similar range. The practical decision is tier positioning: Opus 4 is Anthropic's flagship built for the hardest reasoning, research, and sustained multi-hour agentic tasks, while Sonnet 4 delivers frontier coding performance at substantially lower cost, making it the natural default for high-volume production traffic.

Dimension
Claude 4 Sonnet (Non-reasoning)
Claude 4 Opus (Non-reasoning)
Context window *
200000
tokens
200000
tokens
Output speed *
0.00
tok/s
0.00
tok/s
Intelligence Index *
16.6
16.6
Input pricing
18
AnyToken
90
AnyToken
Output pricing
90
AnyToken
450
AnyToken
Knowledge cutoff *
May 2025
May 2025

Choose Claude Sonnet 4 when you need strong real-world coding and agentic performance at scale, on cost-sensitive, high-throughput workloads such as coding assistants, code review, and customer-facing agents. Choose Claude Opus 4 when a task demands maximum reasoning depth, sustained autonomous operation over many hours, or the most complex research and problem-solving where a capability edge justifies materially higher output costs. Many teams route routine work to Sonnet 4 and escalate only the hardest tasks to Opus 4.

Limitations & Trade-offs

Where Claude 4 Sonnet (Non-reasoning) falls short

1
Not a maximum-reasoning model. Independent Artificial Analysis scoring places Sonnet 4 (non-reasoning) at 26 on its Intelligence Index, above the comparable median but below reasoning-optimized frontier models. For math-heavy, research-grade, or deep multi-step reasoning workloads, a reasoning-first model or Opus-tier model will typically produce better results. Sonnet 4's advantage is concentrated in coding and agentic execution rather than raw reasoning depth.
2
Output length capped at 64K tokens. While the context window reaches 200K (or 1M in beta), maximum output remains 64,000 tokens per response. Workloads that need to generate a single very large artifact, such as an entire long document or a massive code file in one call, must implement chunking or continuation logic. The large input capacity does not translate into proportionally large single-response output.
3
Long-context requests cost more. Prompts exceeding 200,000 tokens are billed at higher long-context input and output rates, and the 1M context remains a public beta feature rather than a default. High-volume applications that routinely push near the 1M limit will see costs rise sharply; prompt caching and batch processing help, but teams should budget carefully before making 1M-token requests a standard part of a pipeline.
4
No enforced JSON output format. Sonnet 4 supports function calling and tool-based structured outputs but does not offer a strict response_format schema enforcement. Applications that require guaranteed, schema-valid JSON must rely on tool-calling patterns and validate responses downstream, adding a small amount of engineering overhead compared with models that enforce structured output natively.

Best-Fit Workloads

Where this model earns its place

01

Autonomous coding agents

‍
Sonnet 4's 72.7% SWE-bench Verified score and near-zero codebase navigation errors make it well suited to agents that read repositories, implement fixes, and pass tests autonomously. Combined with tool calling and extended thinking, it can drive multi-step edit-run-debug loops. Its balanced cost supports the high call volume these agents generate, though the 64K output cap means very large generated files need continuation handling.

02

Code review and bug fixing at scale

‍
For teams running code review, bug triage, and refactoring across large codebases, Sonnet 4 offers frontier coding accuracy at high-volume-friendly cost. The 200K context (or 1M in beta) lets it reason over source files, tests, and documentation together to catch cross-file dependencies. Anthropic positions Sonnet 4 as the balanced default precisely for these recurring engineering tasks rather than one-off flagship-tier problems.

03

Long-document synthesis

‍
The large context window makes Sonnet 4 effective for synthesizing legal contracts, research papers, or technical specifications across many documents in a single request, reducing the need for complex RAG plumbing. Image and PDF input allow it to extract information from charts and scanned pages. For routine document-set analysis this is efficient, but sustained 1M-token usage raises cost and should be reserved for cases that genuinely need it.

04

High-volume customer-facing agents

‍
As Anthropic's balanced tier, Sonnet 4 fits customer support automation and assistant workloads where throughput and cost matter more than maximum reasoning depth. Extended thinking can be disabled for low-latency responses and enabled selectively for harder queries. Its instruction-following and tool-use capabilities support agents that call external APIs, while its predictable cost profile suits sustained production traffic.

Pricing in anytokens via AnyAPI
Input
18
₳
Output
90
₳
Cache write
22.5
₳
Cache read
1.8
₳

Integration

Access Claude 4 Sonnet (Non-reasoning) via AnyAPI.ai

Access Claude 4 Sonnet (Non-reasoning) through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate Claude 4 Sonnet (Non-reasoning) without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access Claude 4 Sonnet (Non-reasoning) and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test Claude 4 Sonnet (Non-reasoning) against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use Claude 4 Sonnet (Non-reasoning) from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use Claude 4 Sonnet (Non-reasoning) for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

Claude Sonnet 4 is best used for real-world software engineering and agentic workflows: autonomous coding agents, code review, bug fixing, and multi-step tool use. It scores 72.7% on SWE-bench Verified and delivers frontier coding performance at Anthropic's balanced-tier cost, making it a strong default for high-volume coding and assistant workloads.

Claude Sonnet 4 has a 200,000-token context window by default, with a 1 million-token context window available in public beta on the Anthropic API, Amazon Bedrock, and Google Cloud Vertex AI. Prompts exceeding 200,000 tokens are billed at higher long-context rates. Maximum output is 64,000 tokens per response, separate from the input context limit.

Yes. Claude Sonnet 4 is a hybrid reasoning model with an optional extended thinking mode that works through problems step by step. Teams can switch between fast standard responses and extended thinking depending on task complexity, trading latency for deeper reasoning only when a request needs it.

Both launched together in May 2025 and share the same architecture, extended thinking, and tool use. Opus 4 is the flagship for maximum reasoning depth and sustained multi-hour agentic tasks; Sonnet 4 delivers frontier coding at substantially lower cost, making it the better default for high-volume production traffic where cost efficiency matters.

Claude Sonnet 4 accepts text and image input, including PDFs, and can extract information from charts and diagrams. It returns text output only; it does not generate images, audio, or video. Its knowledge cutoff is March 2025.

* Benchmark data source: Artificial Analysis artificialanalysis.ai