Mistral AI
•
mistral-nemo
•
Released 
July 2024

Mistral AI
mistral-nemo

A 12B open-weight model with a 128K context window and strong multilingual coverage across eleven languages under Apache 2.0.

Modality:
Text
model ID
mistralai/mistral-nemo

Output Speed *

N/A
tok/s

Intelligence Index *

N/A
/ 100

Context Window *

131072
tokens

Input price

0.12
Anytoken

Output price

0.24
Anytoken
Mistral NeMo: A 128K-Context 12B Model for Multilingual Production Workloads Mistral NeMo is a 12-billion-parameter open-weight text model built by Mistral in collaboration with NVIDIA and released under Apache 2.0. It sits at the compact end of Mistral's lineup as a drop-in replacement for Mistral 7B, pairing a 128,000-token context window with the Tekken tokenizer for efficient multilingual and code compression. Its defining trait is broad language coverage plus a long context at 12B scale, making it well suited to multilingual chat, document processing, and cost-sensitive high-volume inference where a small, permissively licensed model matters. Start calling Mistral NeMo through the AnyAPI.ai unified API.

Performance

Where Mistral NeMo Earns Its Place at 12B Scale

Mistral NeMo's strength is delivering long context and broad multilingual coverage at a small parameter count. It reports 68.0% on MMLU and 83.5% on HellaSwag, competitive figures for its size class where it frequently edges out Gemma 2 9B and Llama 3 8B on commonsense and world-knowledge tasks. Combined with the Tekken tokenizer, which compresses source code and several major languages roughly 30% more efficiently than SentencePiece, the model processes multilingual and code-heavy inputs at lower effective token cost. In production this means economical, single-GPU-deployable multilingual assistants rather than frontier-level reasoning.

Benchmarks

Mistral NeMo Benchmarks: Strong for Size, Not for Frontier Reasoning

On standard evaluations, Mistral NeMo Instruct reports 68.0% MMLU, 83.5% HellaSwag, 76.8% Winogrande, and 73.8% TriviaQA (five-shot). These place it above Llama 3 8B and Mistral 7B on general knowledge and commonsense, while models like Gemma 2 9B and Qwen 2.5 14B score higher on MMLU. It is not designed to compete with dedicated reasoning models such as o1 on math or advanced problem-solving. Interpret the numbers as evidence of a capable general-purpose small model, strongest on world knowledge, multilingual understanding, and long-context tasks rather than complex chained reasoning.

Output Speed

*
N/A
tok/s

Intelligence Index

*
N/A
/ 100

MMLU *

Broad world knowledge and problem-solving
0
%

GPQA *

PhD-level scientific reasoning across physics, biology, chemistry.
0
%

HLE *

Adherence to multi-step structured instructions.
0
%

LiveCodeBench *

Tool-calling reliability in long agentic loops.
0
%

Technical Specifications

What the model supports

Mistral NeMo is a text-in, text-out 12B dense model with a 128,000-token context window and typical hosted output limits around 16,384 tokens. It was trained for function calling and supports JSON output through response_format, though schema enforcement and tool availability vary by endpoint. The most production-relevant characteristics are the long context and the Tekken tokenizer's efficient multilingual and code compression, which lowers effective token consumption for non-English and code-heavy workloads. Note that context window and maximum output are distinct: 128K describes total input capacity, not per-response generation length.
Verified Specifications — 
mistral-nemo
*
Input modalities
Text
output modalities
Text
Context window
131072
 tokens
Maximum output tokens
16384
Reasoning
No
Knowledge cutoff
July 2024
Pricing (standard)
0.12
 AnyTokens in
 / 
0.24
 AnyTokens out

Quickstart

Sample code for mistral-nemo

import requests

url = "https://api.anyapi.ai/v1/chat/completions"

payload = {
    "stream": False,
    "tool_choice": "auto",
    "logprobs": False,
    "model": "Model_Name",
    "messages": [
        {
            "role": "user",
            "content": "Hello"
        }
    ]
}
headers = {
    "Authorization": "Bearer AnyAPI_API_KEY",
    "Content-Type": "application/json"
}

response = requests.post(url, json=payload, headers=headers)

print(response.json())
import requests url = "https://api.anyapi.ai/v1/chat/completions" payload = { "stream": False, "tool_choice": "auto", "logprobs": False, "model": "Model_Name", "messages": [ { "role": "user", "content": "Hello" } ] } headers = { "Authorization": "Bearer AnyAPI_API_KEY", "Content-Type": "application/json" } response = requests.post(url, json=payload, headers=headers) print(response.json())
View docs
Copy
Code is copied
const url = 'https://api.anyapi.ai/v1/chat/completions';
const options = {
  method: 'POST',
  headers: {Authorization: 'Bearer AnyAPI_API_KEY', 'Content-Type': 'application/json'},
  body: '{"stream":false,"tool_choice":"auto","logprobs":false,"model":"Model_Name","messages":[{"role":"user","content":"Hello"}]}'
};

try {
  const response = await fetch(url, options);
  const data = await response.json();
  console.log(data);
} catch (error) {
  console.error(error);
}
const url = 'https://api.anyapi.ai/v1/chat/completions'; const options = { method: 'POST', headers: {Authorization: 'Bearer AnyAPI_API_KEY', 'Content-Type': 'application/json'}, body: '{"stream":false,"tool_choice":"auto","logprobs":false,"model":"Model_Name","messages":[{"role":"user","content":"Hello"}]}' }; try { const response = await fetch(url, options); const data = await response.json(); console.log(data); } catch (error) { console.error(error); }
View docs
Copy
Code is copied
curl --request POST \
  --url https://api.anyapi.ai/v1/chat/completions \
  --header 'Authorization: Bearer AnyAPI_API_KEY' \
  --header 'Content-Type: application/json' \
  --data '{
  "stream": false,
  "tool_choice": "auto",
  "logprobs": false,
  "model": "Model_Name",
  "messages": [
    {
      "role": "user",
      "content": "Hello"
    }
  ]
}'
curl --request POST \ --url https://api.anyapi.ai/v1/chat/completions \ --header 'Authorization: Bearer AnyAPI_API_KEY' \ --header 'Content-Type: application/json' \ --data '{ "stream": false, "tool_choice": "auto", "logprobs": false, "model": "Model_Name", "messages": [ { "role": "user", "content": "Hello" } ] }'
View docs
Copy
Code is copied
View docs
Code examples coming soon...

Limitations & Trade-offs

Where mistral-nemo falls short

1
No native reasoning mode. Mistral NeMo has no dedicated extended-thinking or reasoning-control capability, and its MMLU (68.0%) and general benchmarks trail dedicated reasoning models and larger models like Qwen 2.5 14B. For multi-step math, complex agentic planning, or hard analytical tasks, a purpose-built reasoning model will outperform it. Use NeMo for knowledge, chat, and language tasks rather than heavy reasoning.
2
Text-only modalities. The model accepts text input and produces text output; it has no vision, audio, or document-image capability. Sibling models such as Mistral Small 3.1 add vision, but NeMo does not inherit that. Applications needing image understanding, OCR-style parsing, or multimodal input must route those requests to a different model.
3
General-purpose, not specialist. NeMo is a balanced small model rather than a coding or agent specialist. Dedicated coders like Qwen 2.5 Coder report substantially higher HumanEval scores, and frontier agent models handle complex tool chains better. NeMo's function-calling support also varies by hosting endpoint, so verify tool availability before building agentic workflows on it.
4
Superseded for the 128K niche. Mistral has since released Small 3.1 (24B, 128K context, with vision) as the effective successor to NeMo's long-context positioning, and the original NeMo endpoint is marked deprecated in Mistral's own documentation. It remains widely available through open weights and third-party hosts, but teams starting new projects should weigh newer options for long-lived deployments.

Best-Fit Workloads

Where this model earns its place

01

Multilingual chat assistants

‍
NeMo was designed for global applications with particular strength across eleven languages including French, German, Spanish, Chinese, Japanese, Arabic, and Hindi. The Tekken tokenizer compresses non-English text more efficiently, lowering effective cost per multilingual request. This makes it a strong fit for customer-facing assistants serving diverse language markets where a single small model must handle many languages without per-language routing.

02

Long-document processing

‍
The 128,000-token context window is among the largest in the 12B class, comfortably exceeding the 8K-32K windows of most similarly sized models. This lets you feed long reports, transcripts, or multi-file inputs in a single request for summarization, extraction, or question answering. For non-English or code-heavy documents, Tekken's compression stretches that window further, though NeMo remains a general model rather than a frontier long-context reasoner.

03

Cost-sensitive high-volume inference

‍
At 12B parameters with FP8-ready quantization-aware training, NeMo runs on a single GPU and occupies the lower-cost tier of usable models, making it economical for high-throughput generation. It suits classification, summarization, and templated content at scale where per-request cost matters more than frontier accuracy. Its Apache 2.0 license also permits unrestricted commercial deployment, useful for teams self-hosting to control cost.

04

Multilingual RAG backends

‍
Combining the long context window with broad language coverage makes NeMo a practical generation model for retrieval-augmented pipelines that serve multiple languages. It can absorb sizeable retrieved passages and answer in the user's language. Because its raw reasoning is modest, pair it with strong retrieval and grounding; for complex synthesis over retrieved evidence, a larger model may produce more reliable answers.

Pricing in anytokens via AnyAPI
Input
0.12
₳
Output
0.24
₳
Cache write
—
₳
Cache read
—
₳

Integration

Access mistral-nemo via AnyAPI.ai

Access mistral-nemo through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate mistral-nemo without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access mistral-nemo and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test mistral-nemo against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use mistral-nemo from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use mistral-nemo for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

Mistral NeMo has a 128,000-token context window (commonly reported as 131,072 tokens), one of the largest in its 12B parameter class. Hosted endpoints typically cap output around 16,384 tokens. The large context supports long documents, extended conversations, and multi-file inputs, though total input capacity is separate from per-response generation length.

Yes. Mistral NeMo is released under the Apache 2.0 license, with both pre-trained base and instruction-tuned checkpoints published openly. Apache 2.0 permits unrestricted commercial use, so teams can deploy or self-host the weights freely. It was built by Mistral in collaboration with NVIDIA and named after NVIDIA's NeMo framework.

Mistral NeMo is designed for multilingual use and is particularly strong in English, French, German, Spanish, Italian, Portuguese, Chinese, Japanese, Korean, Arabic, and Hindi. Its Tekken tokenizer was trained on over 100 languages and compresses many of these more efficiently than earlier Mistral tokenizers, improving effective token cost for non-English text.

Mistral NeMo is the intended drop-in replacement for Mistral 7B. It offers a much larger 128K context window, stronger multilingual coverage, and improved reasoning and world-knowledge benchmark scores, at the cost of a larger 12B footprint and higher VRAM requirements. Mistral 7B remains preferable only when memory and throughput are tightly constrained.

Mistral NeMo was trained for function calling and supports JSON output via response_format, but tool availability and schema enforcement vary by hosting endpoint, so verify before building agents. It is text-only and does not support image, audio, or video input. For multimodal needs, use a vision-capable model such as Mistral Small 3.1.

* Benchmark data source: Artificial Analysis artificialanalysis.ai