Cost-efficient open-weight 24B multimodal model tuned for reliable instruction following, tool calling, and fewer repetition failures.
Output Speed *
Intelligence Index *
Context Window *
Input price
Output price
Performance
Where Mistral Small 3.2 Improves Over 3.1
Benchmarks
Mistral Small 3.2 Benchmarks: Strong Formatting, Weak Agentic Depth
Output Speed
Intelligence Index
MMLU *
GPQA *
HLE *
LiveCodeBench *
Technical Specifications
What the model supports
Quickstart
import requests
url = "https://api.anyapi.ai/v1/chat/completions"
payload = {
"stream": False,
"tool_choice": "auto",
"logprobs": False,
"model": "Model_Name",
"messages": [
{
"content": [
{
"type": "text",
"text": "Hello"
},
{
"image_url": {
"detail": "auto",
"url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"
},
"type": "image_url"
}
],
"role": "user"
}
]
}
headers = {
"Authorization": "Bearer AnyAPI_API_KEY",
"Content-Type": "application/json"
}
response = requests.post(url, json=payload, headers=headers)
print(response.json())const url = 'https://api.anyapi.ai/v1/chat/completions';
const options = {
method: 'POST',
headers: {Authorization: 'Bearer AnyAPI_API_KEY', 'Content-Type': 'application/json'},
body: '{"stream":false,"tool_choice":"auto","logprobs":false,"model":"Model_Name","messages":[{"content":[{"type":"text","text":"Hello"},{"image_url":{"detail":"auto","url":"https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"},"type":"image_url"}],"role":"user"}]}'
};
try {
const response = await fetch(url, options);
const data = await response.json();
console.log(data);
} catch (error) {
console.error(error);
}curl --request POST \
--url https://api.anyapi.ai/v1/chat/completions \
--header 'Authorization: Bearer AnyAPI_API_KEY' \
--header 'Content-Type: application/json' \
--data '{
"stream": false,
"tool_choice": "auto",
"logprobs": false,
"model": "Model_Name",
"messages": [
{
"content": [
{
"type": "text",
"text": "Hello"
},
{
"image_url": {
"detail": "auto",
"url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"
},
"type": "image_url"
}
],
"role": "user"
}
]
}'Limitations & Trade-offs
Best-Fit Workloads
Where this model earns its place
Structured Data Extraction
The model's native JSON-schema support plus improved instruction adherence make it well suited to extracting fields from text and documents into strict formats. Reduced repetition and better format compliance mean fewer malformed responses and less retry logic. This benefits invoice parsing, entity extraction, and classification pipelines where output must validate against a schema on the first pass.
Tool-Calling and Function Steps
Mistral upgraded the 3.2 function-calling template specifically for more reliable tool use, particularly via vLLM. This makes it a dependable executor for bounded, well-defined tool-call steps inside larger workflows — API routing, retrieval triggers, and deterministic actions. It is not built for long autonomous agent chains, so keep planning in a reasoning model and delegate individual, specified tool invocations to 3.2.
High-Volume Content and Chat
At roughly 152 tokens/second with sub-second time to first token on Mistral's endpoint, and strong conversational scores on Arena Hard v2 and WildBench v2, the model handles high-throughput chat and content generation efficiently. Its reduced infinite-generation rate improves reliability for user-facing assistants. Best where cost-efficiency and responsiveness matter more than frontier reasoning.
Document and Chart Understanding
With image input and strong DocVQA (~94.9%) and ChartQA (~87.4%) scores, the model reads documents and charts and returns structured text answers. This fits light document analysis, report summarization from scanned pages, and chart-to-text extraction. Note that output is text only and some vision metrics fluctuated versus 3.1, so validate on your specific document types before scaling.