Anthropic
•
Claude 4.5 Haiku (Non-reasoning)
•
Released 
October 2025

Anthropic
Claude 4.5 Haiku (Non-reasoning)

Anthropic's fastest, lowest-cost Claude, delivering Sonnet 4-class coding and agentic performance for high-volume, latency-sensitive production.

Modality:
Text
Image
PDF
model ID
anthropic/claude-haiku-4.5

Output Speed *

101.18
tok/s

Intelligence Index *

15.4
/ 100

Context Window *

200000
tokens

Input price

6.3
Anytoken

Output price

33
Anytoken
Claude Haiku 4.5: Near-Frontier Coding at Haiku Speed and Cost Claude Haiku 4.5 is Anthropic's smallest, fastest, and most cost-efficient Claude model, released October 15, 2025. It sits at the entry tier below Sonnet and Opus, but Anthropic positions it as delivering coding and agentic performance comparable to Claude Sonnet 4—a model considered state of the art months earlier—at roughly one-third the cost and more than twice the speed. It is the first Haiku with an optional extended thinking mode. The workload that benefits most is high-volume, latency-sensitive agentic work: chat assistants, coding sub-agents, classification pipelines, and real-time tool-use loops. Start building with Claude Haiku 4.5 via the AnyAPI.ai API.

Performance

Where Haiku 4.5 Earns Its Place: Speed With Real Coding Ability

Haiku 4.5's defining trait is that it pairs low latency with genuine agentic coding ability rather than trading one for the other. Anthropic reports 73.3% on SWE-bench Verified, comparable to Sonnet 4, while independent testing from Artificial Analysis measures output around 95 tokens per second and time to first token near 0.7 seconds—among the fastest mainstream APIs. In practice this makes tight tool-use feedback loops feel responsive and lets teams run many parallel sub-agents cheaply. The consequence: workloads that were impractical at Sonnet-tier latency and price become viable in production, especially real-time assistants and high-throughput classification.

Benchmarks

Independent Benchmarks: Fast Throughput, Above-Average Intelligence for Its Tier

Artificial Analysis measures Claude Haiku 4.5 output at roughly 95 tokens per second on Anthropic's API—faster than the ~72 t/s median for comparable non-reasoning models—with time to first token near 0.7 seconds, making it a TTFT leader among Anthropic models. On the Artificial Analysis Intelligence Index it scores around 24 (non-reasoning), slightly above the peer median of 23. On coding, Anthropic's reported 73.3% SWE-bench Verified aligns with independent framing that Haiku 4.5 reaches roughly 90% of Sonnet 4.5's agentic coding at a fraction of the cost and several times the speed. Speed and coding, not aggregate reasoning depth, are its measured strengths.

Output Speed

*
101.18
tok/s

Intelligence Index

*
15.4
/ 100

MMLU *

Broad world knowledge and problem-solving
80
%

GPQA *

PhD-level scientific reasoning across physics, biology, chemistry.
65
%

HLE *

Adherence to multi-step structured instructions.
4
%

LiveCodeBench *

Tool-calling reliability in long agentic loops.
51
%

Technical Specifications

What the model supports

Claude Haiku 4.5 accepts text and images (including PDFs) and returns text only. It provides a 200,000-token context window with up to 64,000 output tokens per request—roughly eight times the previous Haiku's output ceiling, which matters for long code generation and detailed rewrites. It is the first Haiku with optional extended thinking, function calling, and structured JSON outputs. The two specifications with the greatest production impact are the 200K context, smaller than the 1M windows on current Sonnet/Opus tiers, and the 64K output, high for a lightweight model. Knowledge cutoff is February 2025.
Verified Specifications — 
Claude 4.5 Haiku (Non-reasoning)
*
Input modalities
Text
Image
PDF
output modalities
Text
Context window
200000
 tokens
Maximum output tokens
64000
Reasoning
Yes
Knowledge cutoff
October 2025
Pricing (standard)
6.3
 AnyTokens in
 / 
33
 AnyTokens out

Comparison

Claude Haiku 4.5 vs Claude Sonnet 4.5: Speed and Cost or Reasoning Depth?

These two are compared because Haiku 4.5 launched two weeks after Sonnet 4.5 and shares the same 200,000-token context window, tool calling, extended thinking, and vision input. Anthropic explicitly positions them to work together: Sonnet 4.5 as the frontier orchestrator and Haiku 4.5 as the fast, cheap sub-agent. The practical decision is not capability parity—Sonnet 4.5 remains the stronger reasoning and coding model—but whether a given task needs that extra depth or benefits more from Haiku's lower latency and substantially lower per-token cost at high volume.

Dimension
Claude 4.5 Haiku (Non-reasoning)
Claude 4.5 Sonnet (Non-reasoning)
Context window *
200000
tokens
1000000
tokens
Output speed *
101.18
tok/s
0.00
tok/s
Intelligence Index *
15.4
19.3
Input pricing
6.3
AnyToken
18
AnyToken
Output pricing
33
AnyToken
90
AnyToken
Knowledge cutoff *
October 2025
September 2025

Choose Claude Haiku 4.5 when latency is user-visible or volume is high: real-time chat, code autocomplete, classification, and fleets of parallel sub-agents where speed and cost dominate. It runs several times faster than Sonnet 4.5 and occupies the lowest-cost tier of the family. Choose Claude Sonnet 4.5 when you need the strongest agentic coding, deeper multi-step reasoning, or long-horizon autonomy on complex tasks, and can absorb higher latency and cost. Many production systems use both—Sonnet to plan, Haiku to execute.

Limitations & Trade-offs

Where Claude 4.5 Haiku (Non-reasoning) falls short

1
Smaller context than current premium tiers. Haiku 4.5 caps input at 200,000 tokens, while current Sonnet and Opus models offer 1M-token windows. For genuinely long single documents—large codebases, 300-page manuals, year-long chat histories—you must chunk and stitch, adding orchestration work and a real risk of missed cross-references between chunks. If your workload centers on single-pass long-context synthesis, a 1M-window model is structurally better suited than Haiku 4.5.
2
Not a frontier reasoning model. Anthropic itself frames Haiku 4.5 as near-frontier, not frontier. Its Artificial Analysis Intelligence Index score (~24 non-reasoning) sits only slightly above the peer median and well below flagship Claude models. For deep multi-layer reasoning, specialized expert synthesis, or the hardest long-horizon agentic tasks, Sonnet or Opus deliver materially higher quality. Haiku's advantage is speed-per-dollar, not maximum intelligence.
3
Extended thinking adds latency and cost. The optional extended thinking mode—Haiku's headline reasoning feature—consumes additional output tokens and increases response time, eroding the speed advantage that is the model's main reason to exist. On latency-sensitive interactive paths, enabling thinking can defeat the purpose. Reserve it for the subset of requests that genuinely need deliberate reasoning, and keep default requests non-reasoning.
4
Text-only output and no verified multilingual claims. Haiku 4.5 accepts text and images but returns text only—no image, audio, or video output. Anthropic does not publish specific multilingual benchmark claims for this model. If your application requires generated media or you need verified performance guarantees in specific non-English languages, validate independently rather than assuming parity with larger or specialist models.

Best-Fit Workloads

Where this model earns its place

01

Real-time chat and customer service agents

‍
Anthropic positions Haiku 4.5 for low-latency assistants, and independent TTFT measurements near 0.7 seconds with ~95 tokens/second output make responses feel near-instant. This suits customer service bots and interactive assistants where every hundred milliseconds is visible to the user. The 200K context comfortably holds conversation history and retrieved documents. The main caveat: for the most complex escalations that need deep reasoning, route to a larger Claude model.

02

Agentic coding and multi-agent sub-agents

‍

With 73.3% on SWE-bench Verified (Anthropic-reported) and coding roughly comparable to Sonnet 4, Haiku 4.5 handles repository-level bug fixing, rapid prototyping, and tool-driven code tasks. Its speed makes tight edit-run-fix loops responsive. Anthropic explicitly recommends it as the sub-agent layer under a Sonnet or Opus orchestrator, so large jobs can be fanned out to many cheap, fast Haiku instances running in parallel.

03

High-volume classification and extraction

‍
The combination of fast throughput, low per-token cost, structured JSON outputs, and function calling makes Haiku 4.5 well suited to mass classification, tagging, moderation, and structured extraction pipelines. Prompt caching further reduces cost when a large shared instruction set or schema is reused across many calls. For these throughput-bound jobs, its speed-per-dollar is the decisive advantage over larger models.

04

Long-output generation and rewrites

‍
The 64,000-token output ceiling—about eight times the previous Haiku generation—lets Haiku 4.5 produce long-form code, detailed reports, and large document rewrites in a single response without stitching multiple calls. Because it proactively shortens output as it approaches the context limit, plan input size so the model has room to complete long generations. For inputs beyond 200K tokens, a 1M-window model remains the better fit.

Pricing in anytokens via AnyAPI
Input
6.3
₳
Output
33
₳
Cache write
—
₳
Cache read
—
₳

Integration

Access Claude 4.5 Haiku (Non-reasoning) via AnyAPI.ai

Access Claude 4.5 Haiku (Non-reasoning) through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate Claude 4.5 Haiku (Non-reasoning) without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access Claude 4.5 Haiku (Non-reasoning) and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test Claude 4.5 Haiku (Non-reasoning) against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use Claude 4.5 Haiku (Non-reasoning) from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use Claude 4.5 Haiku (Non-reasoning) for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

Claude Haiku 4.5 is best for fast, high-volume, latency-sensitive workloads: real-time chat assistants, customer service agents, code autocomplete, classification and extraction pipelines, and sub-agents in multi-agent systems. Anthropic positions it as delivering coding and agentic performance comparable to Sonnet 4 at roughly one-third the cost and more than twice the speed.

Claude Haiku 4.5 has a 200,000-token context window and can generate up to 64,000 output tokens per request. The context window covers all input—system prompts, conversation history, tool outputs, and image tokens—while max output limits how long a single response can be. Current Sonnet and Opus tiers offer larger 1M-token windows.

Yes. Claude Haiku 4.5 is the first Haiku with an optional extended thinking mode for more deliberate reasoning. It also supports function calling via tools and tool_choice, structured JSON outputs, computer use, streaming, and prompt caching. Extended thinking adds latency and token cost, so it is best reserved for tasks that genuinely need deeper reasoning.

Independent testing by Artificial Analysis measures Claude Haiku 4.5 output at roughly 95 tokens per second on Anthropic's API—above the ~72 t/s median for comparable models—with time to first token near 0.7 seconds. That makes it a TTFT leader among Anthropic models. Anthropic states it runs up to 4–5 times faster than Sonnet 4.5.

Use Claude Haiku 4.5 when latency and cost matter at high volume—real-time chat, classification, and parallel sub-agents. Use Sonnet 4.5 when you need the strongest agentic coding, deeper reasoning, or long-horizon autonomy. They share a 200K context and tool support, and Anthropic recommends combining them: Sonnet orchestrates, Haiku executes.

* Benchmark data source: Artificial Analysis artificialanalysis.ai