Meta's open-weight MoE model pairing text-and-image input, a 1M-token window, and fast, low-latency inference for high-volume assistants.
Meta's open-weight MoE model pairing text-and-image input, a 1M-token window, and fast, low-latency inference for high-volume assistants.
Answers to common questions about integrating and using this AI model via AnyAPI.ai
Llama 4 Maverick supports a context window of approximately 1M tokens (1,048,576) per Meta's model card, but its maximum output is capped at 8,192 tokens per response. This makes it ideal for reading and synthesizing very long inputs, but long-form generation must be chunked across multiple requests.
Yes—Maverick accepts both text and image input and returns text output. It is natively multimodal via early fusion, with strong reported results on multimodal benchmarks like MMMU and MathVista. Note that image understanding is English-only, even though its text capabilities span 12 fine-tuned languages.
Independent measurements put median output around 100 tokens per second—well above the roughly 60 t/s median for comparable open-weight non-reasoning models—with sub-second time to first token on optimized FP8 providers. Specialized inference hardware pushes throughput far higher, making Maverick a strong choice for latency-sensitive, high-volume workloads.
Maverick is open-weight under the Llama 4 Community License, which permits commercial use for most organizations but requires a special license above 700 million monthly active users. Importantly, EU-domiciled users and companies are currently restricted from using the models, so confirm license terms before deploying.