Anthropic
•
Claude 4 Opus (Non-reasoning)
•
Released 
May 2025

Anthropic
Claude 4 Opus (Non-reasoning)

Anthropic's launch flagship for long-horizon coding agents and multi-step autonomous workflows that must run for hours.

Modality:
Text
Image
Video
PDF
model ID
anthropic/claude-opus-4

Output Speed *

0.00
tok/s

Intelligence Index *

16.6
/ 100

Context Window *

200000
tokens

Input price

90
Anytoken

Output price

450
Anytoken

Sustained Agentic Coding for Long-Horizon Tasks

‍

Claude Opus 4 is Anthropic's flagship Claude 4 model, released in May 2025 as the most capable model in the family at launch. It is a hybrid reasoning model that switches between near-instant responses and extended thinking, and can interleave reasoning with tool calls in a single workflow. Its defining strength is sustained agentic coding: it maintained multi-hour autonomous task execution and led SWE-bench Verified at release. Teams building coding agents, multi-file refactoring pipelines, and long-running research workflows benefit most from its combination of frontier code performance and stable long-horizon execution.

Access Claude Opus 4 via the AnyAPI.ai API

Performance

Where Claude Opus 4 Earns Its Place: Long-Horizon Coding Agents

Claude Opus 4 was built for complex software engineering and multi-step agentic execution rather than fast, cheap turns. At release it scored 72.5% on SWE-bench Verified (79.4% with parallel test-time compute), leading real-world bug-fixing benchmarks at the time. As a hybrid reasoning model, it interleaves extended thinking with tool calls and sustains coherent multi-hour autonomous work, using local file access to maintain 'memory files' across steps. For production, this means it holds context and intent across long agent trajectories where cheaper models drift, though its Opus-tier economics make it a deliberate choice rather than a high-volume default.

Benchmarks

Claude Opus 4 Benchmarks: Coding, Reasoning and Agentic Scores

On coding, Claude Opus 4 reached 72.5% on SWE-bench Verified at launch, with Terminal-bench at 43.2%. On graduate-level reasoning it scored 79.6% on GPQA Diamond and 75.5% on AIME 2025, and on agentic tool use it reached 81.4% on TAU-bench Retail. Independent measurement from Artificial Analysis places non-reasoning output speed around 32 tokens per second on Anthropic's API with time to first token near 2.1 seconds. The pattern is clear: frontier accuracy and agentic reliability paired with modest throughput, favoring correctness-critical workloads over latency-sensitive ones.

Output Speed

*
0.00
tok/s

Intelligence Index

*
16.6
/ 100

MMLU *

Broad world knowledge and problem-solving
86
%

GPQA *

PhD-level scientific reasoning across physics, biology, chemistry.
70
%

HLE *

Adherence to multi-step structured instructions.
6
%

LiveCodeBench *

Tool-calling reliability in long agentic loops.
54
%

Technical Specifications

What the model supports

Claude Opus 4 accepts text and images (including PDFs) and returns text; it is not an image, audio, or video generator. It provides a 200,000-token context window with up to 32,000 output tokens per response, so the context window is working memory for long inputs while output length is separately capped. Extended thinking, function/tool calling, streaming, prompt caching, and structured outputs are supported. The 32K output ceiling means very large single-shot generations must be chunked, and the 200K window is smaller than the 1M window Anthropic later shipped on newer Opus versions.
Verified Specifications — 
Claude 4 Opus (Non-reasoning)
*
Input modalities
Text
Image
Video
PDF
output modalities
Text
Context window
200000
 tokens
Maximum output tokens
32000
Reasoning
Yes
Knowledge cutoff
May 2025
Pricing (standard)
90
 AnyTokens in
 / 
450
 AnyTokens out

Quickstart

Sample code for Claude 4 Opus (Non-reasoning)

import requests

url = "https://api.anyapi.ai/v1/chat/completions"

payload = {
    "model": "claude-4-opus",
    "messages": [
        {
            "role": "user",
            "content": "Hello"
        }
    ]
}
headers = {
    "Authorization": "Bearer  AnyAPI_API_KEY",
    "Content-Type": "application/json"
}

response = requests.post(url, json=payload, headers=headers)

print(response.json())
import requests url = "https://api.anyapi.ai/v1/chat/completions" payload = { "model": "claude-4-opus", "messages": [ { "role": "user", "content": "Hello" } ] } headers = { "Authorization": "Bearer AnyAPI_API_KEY", "Content-Type": "application/json" } response = requests.post(url, json=payload, headers=headers) print(response.json())
View docs
Copy
Code is copied
const url = 'https://api.anyapi.ai/v1/chat/completions';
const options = {
  method: 'POST',
  headers: {Authorization: 'Bearer  AnyAPI_API_KEY', 'Content-Type': 'application/json'},
  body: '{"model":"claude-4-opus","messages":[{"role":"user","content":"Hello"}]}'
};

try {
  const response = await fetch(url, options);
  const data = await response.json();
  console.log(data);
} catch (error) {
  console.error(error);
}
const url = 'https://api.anyapi.ai/v1/chat/completions'; const options = { method: 'POST', headers: {Authorization: 'Bearer AnyAPI_API_KEY', 'Content-Type': 'application/json'}, body: '{"model":"claude-4-opus","messages":[{"role":"user","content":"Hello"}]}' }; try { const response = await fetch(url, options); const data = await response.json(); console.log(data); } catch (error) { console.error(error); }
View docs
Copy
Code is copied
curl --request POST \
  --url https://api.anyapi.ai/v1/chat/completions \
  --header 'Authorization: Bearer  AnyAPI_API_KEY' \
  --header 'Content-Type: application/json' \
  --data '{
  "model": "claude-4-opus",
  "messages": [
    {
      "role": "user",
      "content": "Hello"
    }
  ]
}'
curl --request POST \ --url https://api.anyapi.ai/v1/chat/completions \ --header 'Authorization: Bearer AnyAPI_API_KEY' \ --header 'Content-Type: application/json' \ --data '{ "model": "claude-4-opus", "messages": [ { "role": "user", "content": "Hello" } ] }'
View docs
Copy
Code is copied
View docs
Code examples coming soon...

Comparison

Claude Opus 4 vs Claude Opus 4.1: What Actually Changes?

Claude Opus 4.1 is the direct successor in the same Opus tier and shares the 200,000-token context window, 32,000-token output cap, extended thinking, and the same input modalities. The realistic decision is whether the incremental coding and agentic gains of 4.1 justify moving off Opus 4. Opus 4.1 raised SWE-bench Verified to 74.5% and delivered measurable improvements on autonomous coding trajectories, while keeping the same API surface. Because both occupy the premium Opus tier, this is not a cost-versus-capability trade so much as a straightforward capability refresh within one model family.

Dimension
Claude 4 Opus (Non-reasoning)
Claude 4.1 Opus (Non-reasoning)
Context window *
200000
tokens
200000
tokens
Output speed *
0.00
tok/s
0.00
tok/s
Intelligence Index *
16.6
18.6
Input pricing
90
AnyToken
90
AnyToken
Output pricing
450
AnyToken
450
AnyToken
Knowledge cutoff *
May 2025
August 2025

Choose Claude Opus 4 when you have already validated and pinned it in production and need behavioral stability, or when you specifically require the original Opus 4 snapshot for reproducible evaluations. Choose Claude Opus 4.1 when you want the stronger coding and agentic scores on the same API with no architectural changes, which is the sensible default for new integrations that would otherwise reach for Opus 4. Neither changes context or output limits, so migration is low-risk.

Limitations & Trade-offs

Where Claude 4 Opus (Non-reasoning) falls short

1
Modest output speed. Independent Artificial Analysis measurement puts Claude Opus 4 around 32 tokens per second on Anthropic's API with time to first token near 2.1 seconds. That throughput is fine for background agents and batch work but ill-suited to latency-sensitive, interactive chat where users expect fast streaming. For real-time assistants, a Sonnet- or Haiku-tier model is a better fit.
2
Premium Opus-tier economics. Claude Opus 4 sits in Anthropic's highest-cost tier, and long outputs are comparatively expensive. This makes it a deliberate choice for correctness-critical coding and agentic tasks rather than a high-volume default. Prompt caching and batch processing reduce cost meaningfully, but for large-scale generation or high-QPS chat, a cheaper model is the economical option.
3
32K output ceiling and 200K context. Maximum output is capped at 32,000 tokens per response and the context window is 200,000 tokens. Very large single-shot generations must be chunked, and workloads that need to hold an entire large monorepo or a 1M-token corpus in one call exceed this window. Later Opus versions ship a 1M-token window; if you need that capacity, Opus 4 is not the right version.
4
Superseded as the recommended Opus tier. Anthropic has since released Opus 4.1 and later versions that improve on coding, agentic, and reasoning benchmarks. Opus 4 remains a stable, pinned snapshot, which is valuable for reproducibility, but new projects choosing purely on capability will generally get more from a newer Opus release on the same API.

Best-Fit Workloads

Where this model earns its place

01

Autonomous coding agents

‍
Claude Opus 4 was designed for sustained, multi-hour agentic coding and can run background tasks that span thousands of steps. Its 72.5% SWE-bench Verified score and ability to maintain 'memory files' across steps make it well-suited to agents that resolve real issues, refactor across files, and self-correct over long trajectories. The main constraint is throughput and cost, so use it where correctness matters more than speed.

02

Large-scale multi-file refactoring

‍
The combination of frontier code accuracy and the 200,000-token context window suits refactoring tasks that touch many files and require reasoning over a substantial slice of a codebase at once. Extended thinking helps the model plan changes before editing. Be mindful of the 32K output cap for very large diffs and consider chunking generation across multiple calls.

03

Long-document research synthesis

‍
With a 200K-token window and strong reasoning scores (79.6% GPQA Diamond), Opus 4 can synthesize findings across many documents in a single session and interleave web search or retrieval tool calls. This fits deep-research assistants and analyst workflows where depth and accuracy outrank latency. For corpora beyond 200K tokens you will need retrieval or a larger-context model.

04

Multi-step tool-using workflows

‍
Opus 4's hybrid reasoning and function calling let it plan, call tools, observe results, and continue reasoning within one workflow, reaching 81.4% on TAU-bench Retail for multi-step task completion. This suits orchestration and back-office automation that chain several tools reliably. The trade-off is that deliberate multi-step responses cost more and run slower than lightweight models.

Pricing in anytokens via AnyAPI
Input
90
₳
Output
450
₳
Cache write
112.5
₳
Cache read
9
₳

Integration

Access Claude 4 Opus (Non-reasoning) via AnyAPI.ai

Access Claude 4 Opus (Non-reasoning) through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate Claude 4 Opus (Non-reasoning) without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access Claude 4 Opus (Non-reasoning) and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test Claude 4 Opus (Non-reasoning) against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use Claude 4 Opus (Non-reasoning) from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use Claude 4 Opus (Non-reasoning) for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

Claude Opus 4 has a 200,000-token context window and can generate up to 32,000 output tokens per response via the Messages API. The context window is total working memory for input plus output, while the 32,000-token limit caps a single response. For larger single-shot generations, chunk the output across multiple calls.

Yes. Claude Opus 4 was benchmarked as a leading coding model at its May 2025 release, scoring 72.5% on SWE-bench Verified (79.4% with parallel test-time compute) and 43.2% on Terminal-bench. It excels at multi-file refactoring, resolving real software issues, and sustained autonomous coding agents that run over long, multi-step trajectories.

Yes. Claude Opus 4 is a hybrid reasoning model that switches between near-instant responses and extended thinking with a developer-controllable thinking budget. It supports function calling via tools and tool_choice, streaming, prompt caching, and structured outputs, and can interleave reasoning with tool calls such as web search inside a single workflow.

Yes. Claude Opus 4 accepts text, images, and documents such as PDFs as input, and returns text output. It does not generate images, audio, or video. This makes it suitable for tasks that combine visual context with reasoning, but its output modality is text only.

Opus 4.1 is the direct successor and shares Opus 4's 200,000-token context, 32,000-token output cap, and API surface, while improving coding and agentic performance (74.5% SWE-bench Verified versus 72.5%). Choose Opus 4 for a pinned, stable snapshot; choose Opus 4.1 for stronger capability on the same integration.

* Benchmark data source: Artificial Analysis artificialanalysis.ai