Nous Research
•
Hermes 4 - Llama-3.1 405B (Reasoning)
•
Released 
August 2025

Nous Research
Hermes 4 - Llama-3.1 405B (Reasoning)

Open-weight 405B hybrid-reasoning model with a single-checkpoint think toggle and deliberately neutral, low-refusal alignment.

Modality:
Text
PDF
model ID
nousresearch/hermes-4-405b

Output Speed *

39.98
tok/s

Intelligence Index *

7.5
/ 100

Context Window *

131072
tokens

Input price

6
Anytoken

Output price

18
Anytoken
Hermes 4 405B: One Open-Weight Checkpoint That Reasons or Responds on Demand Hermes 4 405B is an open-weight, hybrid-reasoning model from Nous Research, built as a full fine-tune of Meta's Llama-3.1-405B. It is the flagship of the Hermes 4 line, sitting above the 70B variant. Its defining trait is a single checkpoint that switches between direct answering and explicit chain-of-thought via a toggle, controlled by a reasoning parameter. Trained on a ~60B-token corpus emphasizing verified reasoning traces, it targets math, code, STEM, and logic workloads while maintaining a neutral, low-refusal alignment stance that suits teams needing steerable, uncensored behavior. Start calling Hermes 4 405B through the AnyAPI.ai unified API today.

Performance

Hermes 4 405B: Reasoning-On Math and Logic vs Real-World Throughput

Hermes 4 405B is strongest on structured reasoning when its think mode is enabled. The technical report reports 96.3% on MATH-500, 81.9% on AIME 2024, and 61.3% on LiveCodeBench, competitive with other open-weight reasoning models. Those figures are the reasoning-on numbers, which many hosted deployments do not enable by default. This matters because the same weights behave very differently depending on invocation: leave reasoning off and you trade deliberation for latency. In production, teams must explicitly manage the toggle to get the model's headline math and logic performance rather than its faster direct-answer mode.

Benchmarks

How Hermes 4 405B Scores on Independent Evaluations

Independent testing tells a more cautious story than the technical report. Artificial Analysis benchmarks the widely hosted non-reasoning configuration and places Hermes 4 405B below the median on its composite Intelligence Index, ranking it in the lower portion of tracked models. Measured output runs around 36–37 tokens per second with a time to first token near 2.38 seconds, both slower than the median for comparable open-weight models. The gap reflects benchmark conditions: vendor-reported scores use reasoning-on settings, while independent measurement of the default non-reasoning endpoint understates the model's math and logic ceiling.

Output Speed

*
39.98
tok/s

Intelligence Index

*
7.5
/ 100

MMLU *

Broad world knowledge and problem-solving
83
%

GPQA *

PhD-level scientific reasoning across physics, biology, chemistry.
73
%

HLE *

Adherence to multi-step structured instructions.
11
%

LiveCodeBench *

Tool-calling reliability in long agentic loops.
69
%

Technical Specifications

What the model supports

Hermes 4 405B is a dense, text-only model with a 131,072-token context window and text output only; it is not multimodal. Its production-defining specification is the hybrid reasoning toggle, controlled via a reasoning boolean, letting one checkpoint emit explicit <think> traces or answer directly. Open weights (BF16, FP8, GGUF) under the Llama 3.1 Community License permit self-hosting and commercial use with restrictions. Note that some hosted endpoints expose JSON output via response_format without strict schema enforcement, and tool availability varies by provider, so verify function-calling support on your chosen endpoint.
Verified Specifications — 
Hermes 4 - Llama-3.1 405B (Reasoning)
*
Input modalities
Text
PDF
output modalities
Text
Context window
131072
 tokens
Maximum output tokens
117964
Reasoning
Yes
Knowledge cutoff
August 2025
Pricing (standard)
6
 AnyTokens in
 / 
18
 AnyTokens out

Quickstart

Sample code for Hermes 4 - Llama-3.1 405B (Reasoning)

import requests

url = "https://api.anyapi.ai/v1/chat/completions"

payload = {
    "stream": False,
    "tool_choice": "auto",
    "logprobs": False,
    "model": "Model_Name",
    "messages": []
}
headers = {
    "Authorization": "Bearer AnyAPI_API_KEY",
    "Content-Type": "application/json"
}

response = requests.post(url, json=payload, headers=headers)

print(response.json())
import requests url = "https://api.anyapi.ai/v1/chat/completions" payload = { "stream": False, "tool_choice": "auto", "logprobs": False, "model": "Model_Name", "messages": [] } headers = { "Authorization": "Bearer AnyAPI_API_KEY", "Content-Type": "application/json" } response = requests.post(url, json=payload, headers=headers) print(response.json())
View docs
Copy
Code is copied
const url = 'https://api.anyapi.ai/v1/chat/completions';
const options = {
  method: 'POST',
  headers: {Authorization: 'Bearer AnyAPI_API_KEY', 'Content-Type': 'application/json'},
  body: '{"stream":false,"tool_choice":"auto","logprobs":false,"model":"Model_Name","messages":[]}'
};

try {
  const response = await fetch(url, options);
  const data = await response.json();
  console.log(data);
} catch (error) {
  console.error(error);
}
const url = 'https://api.anyapi.ai/v1/chat/completions'; const options = { method: 'POST', headers: {Authorization: 'Bearer AnyAPI_API_KEY', 'Content-Type': 'application/json'}, body: '{"stream":false,"tool_choice":"auto","logprobs":false,"model":"Model_Name","messages":[]}' }; try { const response = await fetch(url, options); const data = await response.json(); console.log(data); } catch (error) { console.error(error); }
View docs
Copy
Code is copied
curl --request POST \
  --url https://api.anyapi.ai/v1/chat/completions \
  --header 'Authorization: Bearer AnyAPI_API_KEY' \
  --header 'Content-Type: application/json' \
  --data '{
  "stream": false,
  "tool_choice": "auto",
  "logprobs": false,
  "model": "Model_Name",
  "messages": []
}'
curl --request POST \ --url https://api.anyapi.ai/v1/chat/completions \ --header 'Authorization: Bearer AnyAPI_API_KEY' \ --header 'Content-Type: application/json' \ --data '{ "stream": false, "tool_choice": "auto", "logprobs": false, "model": "Model_Name", "messages": [] }'
View docs
Copy
Code is copied
View docs
Code examples coming soon...

Comparison

Hermes 4 405B vs Hermes 4 70B: Which Hybrid-Reasoning Checkpoint?

Both models share the same Hermes 4 design: a single hybrid-reasoning checkpoint with a <think> toggle, structured outputs, and Nous Research's neutral, low-refusal alignment. Both are open-weight fine-tunes of Llama 3.1. The practical difference is scale. The 405B variant is a full fine-tune of Llama-3.1-405B and targets the strongest math, code, and logic ceiling, while the 70B variant offers the same behaviors on a far smaller, cheaper-to-serve base. The decision is a familiar one: maximum capability at high compute cost versus a lighter footprint that is easier to self-host and run at volume.

Dimension
Hermes 4 - Llama-3.1 405B (Reasoning)
Hermes 4 - Llama-3.1 70B (Reasoning)
Context window *
131072
tokens
131072
tokens
Output speed *
39.98
tok/s
0.00
tok/s
Intelligence Index *
7.5
7.9
Input pricing
6
AnyToken
0.78
AnyToken
Output pricing
18
AnyToken
2.4
AnyToken
Knowledge cutoff *
August 2025
August 2025

Choose Hermes 4 405B when you need the family's highest reasoning ceiling for hard math, logic, and code, and can absorb slow output speed and higher serving cost. Choose Hermes 4 70B when latency, throughput, and self-hosting economics matter more than the last increment of accuracy, or when you plan to run the model on more modest GPU infrastructure. For high-volume assistant traffic, the 70B is the more pragmatic default; reserve the 405B for the hardest deliberative tasks.

Limitations & Trade-offs

Where Hermes 4 - Llama-3.1 405B (Reasoning) falls short

1
Slow output and higher latency. Independent measurement puts Hermes 4 405B around 36–37 tokens per second with a time to first token near 2.38 seconds, both below the median for comparable open-weight models. As a dense 406B model, generation is inherently heavy, and reasoning-on mode adds further token overhead. This makes it a poor fit for real-time chat or interactive low-latency UX. For latency-sensitive applications, a smaller model or a faster MoE alternative will serve users noticeably better.
2
Below-median independent intelligence scores. Artificial Analysis ranks the widely hosted non-reasoning configuration below the median on its composite Intelligence Index. Part of this is a benchmark-conditions artifact — vendor math/logic numbers use reasoning-on settings — but developers relying on default endpoints without enabling the toggle should not expect frontier general-purpose intelligence. For broad knowledge tasks or agentic benchmarks, newer models often score materially higher.
3
No multimodality. Hermes 4 405B accepts text input and produces text output only. There is no image, audio, video, or document-vision capability. Any workload requiring visual understanding, chart parsing, or audio transcription needs a different model. This restricts the 405B to purely textual reasoning, code, and language tasks.
4
Neutral, low-refusal alignment shifts safety burden to you. Nous reports state-of-the-art RefusalBench results reflecting a deliberate 'aligned to you' stance with low refusal rates. That flexibility is valuable for steerable, uncensored applications, but it means the model applies fewer built-in guardrails than heavily-aligned proprietary models. Teams deploying it in consumer-facing or regulated contexts must add their own content moderation and safety layers.

Best-Fit Workloads

Where this model earns its place

01

Math, logic and STEM problem solving

‍
With reasoning enabled, Hermes 4 405B reports 96.3% on MATH-500 and 81.9% on AIME 2024 in the technical report, making it a credible open-weight choice for deliberative quantitative work. The explicit  traces are inspectable, which helps debugging and verification. Best applied to batch or asynchronous problem-solving where the slow output speed is acceptable and correctness matters more than response time.

02

Code generation and technical reasoning

‍
Hermes 4 405B reports 61.3% on LiveCodeBench with reasoning on, and its post-training corpus heavily emphasizes code and logic. Combined with structured-output support, it fits code-generation pipelines that can tolerate slower generation. Suited to offline code synthesis, refactoring proposals, and technical reasoning tasks rather than latency-critical inline completion, where faster specialized coding models are preferable.

03

Steerable, low-refusal assistants and roleplay

‍
The model's neutral alignment and low refusal rates make it well-suited to applications requiring strong persona adherence, uncensored creative writing, or user-directed behavior that heavily-aligned models often decline. Open weights allow full control and self-hosting. Teams must supply their own moderation layer, but for research, creative tooling, and character-driven products this steerability is a genuine differentiator.

04

Long-document analysis and RAG synthesis

‍
The 131K-token context window lets Hermes 4 405B ingest large documents and long conversation histories in a single request, useful for RAG and multi-document synthesis. Its reasoning mode supports careful cross-referencing across retrieved context. Because output is slow, it fits batch document analysis and summarization pipelines rather than interactive querying, and lacks native vision for scanned or image-based documents.

Pricing in anytokens via AnyAPI
Input
6
₳
Output
18
₳
Cache write
—
₳
Cache read
—
₳

Integration

Access Hermes 4 - Llama-3.1 405B (Reasoning) via AnyAPI.ai

Access Hermes 4 - Llama-3.1 405B (Reasoning) through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate Hermes 4 - Llama-3.1 405B (Reasoning) without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access Hermes 4 - Llama-3.1 405B (Reasoning) and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test Hermes 4 - Llama-3.1 405B (Reasoning) against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use Hermes 4 - Llama-3.1 405B (Reasoning) from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use Hermes 4 - Llama-3.1 405B (Reasoning) for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

Hermes 4 405B has a 131,072-token context window, shared across your prompt and conversation history. This supports long documents, extended multi-turn sessions, and RAG pipelines. It is a text-only window; the model does not accept image, audio, or video inputs.

Hermes 4 405B is a single checkpoint with a hybrid reasoning mode. Setting the reasoning parameter to enabled makes the model emit explicit <think>...</think> traces before answering; disabling it gives faster direct responses. The vendor's headline math and code scores are the reasoning-on numbers, which many hosted endpoints do not enable by default.

Yes. Hermes 4 405B is released with open weights in BF16, FP8, and GGUF formats under the Meta Llama 3.1 Community License. It can be downloaded and self-hosted, and commercial use is permitted with the license's restrictions. It is a 406-billion-parameter dense fine-tune of Llama-3.1-405B.

Independent measurement by Artificial Analysis reports roughly 36–37 output tokens per second and a time to first token near 2.38 seconds for the non-reasoning configuration, both slower than the median for comparable open-weight models. It is not well suited to real-time, latency-sensitive applications.

It is best for deliberative reasoning tasks — math, logic, STEM, and code generation — where correctness matters more than speed, and for steerable, low-refusal assistant and creative applications. Its slow output and text-only design make it a poor fit for real-time chat or multimodal workloads.

* Benchmark data source: Artificial Analysis artificialanalysis.ai