Meta Llama
•
Llama 4 Maverick
•
Released 
April 2025

Meta Llama
Llama 4 Maverick

Meta's open-weight MoE model pairing text-and-image input, a 1M-token window, and fast, low-latency inference for high-volume assistants.

Modality:
Text
Image
model ID
meta-llama/llama-4-maverick

Output Speed *

68.64
tok/s

Intelligence Index *

10
/ 100

Context Window *

1048576
tokens

Input price

0.9
Anytoken

Output price

3.6
Anytoken
Llama 4 Maverick: Fast, Open-Weight Multimodal Inference at Scale Llama 4 Maverick is Meta's larger Llama 4 model, a natively multimodal mixture-of-experts system with 17B active parameters drawn from 400B total across 128 experts. It sits above Llama 4 Scout in capability while sharing the same active-parameter count, so inference stays fast. Maverick accepts text and image input, returns text, and supports a 1M-token context window. Its strongest fit is high-throughput, cost-sensitive production workloads—multilingual assistants, document and image understanding, and streaming chat—where fast output and low first-token latency matter more than frontier reasoning. Start calling Llama 4 Maverick through the AnyAPI.ai API

Performance

Where Maverick Earns Its Place: Speed and Multimodal Throughput

Maverick's defining strength is fast, low-latency multimodal serving rather than frontier intelligence. Independent measurements place median output around 100 tokens per second—well above the roughly 60 t/s median for comparable open-weight non-reasoning models—with sub-second time to first token on optimized FP8 providers. Because the MoE architecture activates only 17B parameters per token, throughput stays high despite the 400B total. For production, this means Maverick suits streaming chat, high-volume assistants, and image-plus-text pipelines where responsiveness and cost efficiency dominate, and where you don't require top-tier reasoning or agentic accuracy.

Benchmarks

Llama 4 Maverick Benchmarks: Fast and Efficient, Not Frontier

Independent testing tells a consistent story. Artificial Analysis scores Maverick below the median on its Intelligence Index while rating it notably fast and reasonably priced among open-weight non-reasoning peers. Meta's reported model-card figures show strong multimodal results (MMMU 73.4, MathVista 73.7) and competitive MMLU Pro, but agentic and novel-reasoning tests are weak—Maverick scored near zero on ARC-AGI-2 in independent runs, and coding percentiles sit low. Note that Meta's high LMArena ELO came from an experimental, style-tuned variant; the public model ranked far lower. Evaluate on your own tasks.

Output Speed

*
68.64
tok/s

Intelligence Index

*
10
/ 100

MMLU *

Broad world knowledge and problem-solving
81
%

GPQA *

PhD-level scientific reasoning across physics, biology, chemistry.
67
%

HLE *

Adherence to multi-step structured instructions.
5
%

LiveCodeBench *

Tool-calling reliability in long agentic loops.
40
%

Technical Specifications

What the model supports

Maverick accepts text and image input and returns text only; image understanding is English-only, even though text spans 12 fine-tuned languages. The 1M-token context window is large, but maximum output is capped at 8,192 tokens, so it fits long-context reading far better than long-form generation. It is a non-reasoning model with no extended-thinking mode. Tool calling and JSON-schema structured outputs are supported. Note the knowledge cutoff of August 2024 and the license restriction prohibiting EU-domiciled use, which affects deployment planning.
Verified Specifications — 
Llama 4 Maverick
*
Input modalities
Text
Image
output modalities
Text
Context window
1048576
 tokens
Maximum output tokens
16384
Reasoning
No
Knowledge cutoff
April 2025
Pricing (standard)
0.9
 AnyTokens in
 / 
3.6
 AnyTokens out

Limitations & Trade-offs

Where Llama 4 Maverick falls short

1
Below-average intelligence for demanding tasks. Independent evaluation places Maverick under the median on the Artificial Analysis Intelligence Index and low on coding and math percentiles. For complex reasoning, agentic planning, or hard coding, frontier or reasoning-capable models will outperform it. Maverick is a fast, general-purpose workhorse, not a reasoning specialist—route difficult branches elsewhere.
2
No reasoning mode and weak agentic scores. Maverick has no extended-thinking configuration, and independent agentic benchmarks (including near-zero ARC-AGI-2 results) show it is a poor fit for autonomous multi-step agents. If your workload depends on reliable tool-chaining, long planning horizons, or self-correction, a dedicated reasoning or agentic model is preferable despite Maverick's speed and cost advantages.
3
Short maximum output despite a huge context window. The 1M-token context is paired with an 8,192-token output cap, so Maverick reads far more than it can write in one turn. Long-form generation—large reports, extensive code files, book-length drafts—requires chunking or a different model. Treat the large window as an input/retrieval advantage, not a long-output capability.
4
Deployment and licensing constraints. Self-hosting Maverick realistically needs a full H100 DGX-class host (FP8) rather than a single consumer GPU, and the Llama 4 Community License restricts EU-domiciled use and imposes a 700M-monthly-active-user threshold. Teams in the EU, or those without enterprise GPU access, should plan around hosted APIs and confirm license terms before production.

Best-Fit Workloads

Where this model earns its place

01
High-volume multilingual chat assistants Maverick's fast output (~100 t/s median) and low first-token latency make it well suited to streaming, interactive assistants at scale. With 12 fine-tuned languages and strong steerability via system prompts, it handles multilingual customer-facing chat cost-efficiently. Keep responses within the 8,192-token cap and route genuinely hard reasoning to a stronger model.
02
Image and document understanding Native text-and-image input supports visual Q&A, captioning, chart and document analysis, with strong reported multimodal benchmarks (MMMU 73.4, MathVista 73.7). Combined with the large context window, Maverick fits pipelines that ingest images alongside long text. Note image understanding is English-only, which constrains non-English visual workloads.
03
Long-context retrieval and RAG The ~1M-token window lets Maverick ingest large documents, transcripts, or retrieved passages in a single request, while fast inference keeps RAG responses snappy. It excels at reading and synthesizing long inputs rather than producing long outputs, so pair it with retrieval and keep generations concise. Independent testing suggests validating long-context accuracy on your own data.
04
Latency-sensitive prefill and first-turn responses With the lowest measured time to first token in several independent comparisons, Maverick is a strong choice for autocomplete, predictive typing, and warm-up turns where perceived responsiveness dominates. In multi-model routes it can take the first turn or easy branches while smarter, slower models handle difficult reasoning steps.
Pricing in anytokens via AnyAPI
Input
0.9
₳
Output
3.6
₳
Cache write
—
₳
Cache read
—
₳

Integration

Access Llama 4 Maverick via AnyAPI.ai

Access Llama 4 Maverick through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate Llama 4 Maverick without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access Llama 4 Maverick and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test Llama 4 Maverick against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use Llama 4 Maverick from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use Llama 4 Maverick for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

Llama 4 Maverick supports a context window of approximately 1M tokens (1,048,576) per Meta's model card, but its maximum output is capped at 8,192 tokens per response. This makes it ideal for reading and synthesizing very long inputs, but long-form generation must be chunked across multiple requests.

Yes—Maverick accepts both text and image input and returns text output. It is natively multimodal via early fusion, with strong reported results on multimodal benchmarks like MMMU and MathVista. Note that image understanding is English-only, even though its text capabilities span 12 fine-tuned languages.

No. Maverick is a non-reasoning model with no extended-thinking configuration. Independent benchmarks show below-median intelligence and weak agentic scores, so for complex reasoning, planning, or autonomous agents a dedicated reasoning model is a better fit. Maverick's advantage is speed, multimodality, and cost efficiency.

Independent measurements put median output around 100 tokens per second—well above the roughly 60 t/s median for comparable open-weight non-reasoning models—with sub-second time to first token on optimized FP8 providers. Specialized inference hardware pushes throughput far higher, making Maverick a strong choice for latency-sensitive, high-volume workloads.

Maverick is open-weight under the Llama 4 Community License, which permits commercial use for most organizations but requires a special license above 700 million monthly active users. Importantly, EU-domiciled users and companies are currently restricted from using the models, so confirm license terms before deploying.

* Benchmark data source: Artificial Analysis artificialanalysis.ai