Mistral AI
•
Mistral Small 3.2
•
Released 
June 2025

Mistral AI
Mistral Small 3.2

Cost-efficient open-weight 24B multimodal model tuned for reliable instruction following, tool calling, and fewer repetition failures.

Modality:
Text
Image
model ID
mistralai/mistral-small-3.2-24b-instruct

Output Speed *

N/A
tok/s

Intelligence Index *

8.2
/ 100

Context Window *

128000
tokens

Input price

0.45
Anytoken

Output price

1.2
Anytoken
Mistral Small 3.2 24B: Reliable Instruction Following at High Throughput Mistral Small 3.2 24B is an Apache 2.0 open-weight, 24-billion-parameter model from Mistral AI and an incremental update to Mistral Small 3.1. It keeps the same architecture, vision input, and 128K context while targeting three concrete weaknesses: instruction adherence, output stability, and function-calling robustness. It sits in the balanced, cost-efficient tier of Mistral's lineup rather than the flagship reasoning tier. It fits high-volume workloads that need dependable structured outputs, tool use, and low latency, where predictable formatting matters more than frontier reasoning depth. Start building with the Mistral Small 3.2 24B API on AnyAPI.ai

Performance

Where Mistral Small 3.2 Improves Over 3.1

Mistral Small 3.2 is tuned for reliability rather than raw capability gains. The most meaningful improvement is output stability: Mistral reports infinite or repetitive generations dropping from 2.11% to 1.29%, roughly halving a failure mode that breaks production pipelines. Instruction-following accuracy on Mistral's internal WildBench rose from about 82.8% to 84.8%. For developers, this matters because bounded, correctly-formatted responses reduce retry logic and post-processing. Combined with fast output (independently measured near 152 tokens/second), the model suits high-volume generation where predictable behavior and throughput outweigh frontier reasoning depth.

Benchmarks

Mistral Small 3.2 Benchmarks: Strong Formatting, Weak Agentic Depth

Independent testing frames the model as capable but not agentic. Artificial Analysis assigns an Intelligence Index of 11, above the ~6 median for comparable open-weight non-reasoning models, at roughly 152 tokens/second. Coding scores are solid for the size: Mistral reports HumanEval Plus at 92.9% and MBPP Plus at 78.33%. Conversational benchmarks jumped notably over 3.1, with Arena Hard v2 more than doubling to ~43.1% and WildBench v2 near 65%. However, agentic and terminal benchmarks (Tau2-Bench, Terminal-Bench, SciCode) rank low, confirming this is not a model for autonomous multi-step agents.

Output Speed

*
N/A
tok/s

Intelligence Index

*
8.2
/ 100

MMLU *

Broad world knowledge and problem-solving
68
%

GPQA *

PhD-level scientific reasoning across physics, biology, chemistry.
51
%

HLE *

Adherence to multi-step structured instructions.
4
%

LiveCodeBench *

Tool-calling reliability in long agentic loops.
28
%

Technical Specifications

What the model supports

Mistral Small 3.2 24B accepts text and image input and returns text only. It carries a 128K-token context window, with hosted providers commonly capping single-request output around 16K tokens. The two production-relevant characteristics are its improved function-calling template (more reliable tool use, especially via vLLM) and native JSON/structured output support. It is a standard instruct model with no dedicated extended-thinking mode, so reasoning-heavy chains should be handled by a reasoning-tier model instead.
Verified Specifications — 
Mistral Small 3.2
*
Input modalities
Text
Image
output modalities
Text
Context window
128000
 tokens
Maximum output tokens
16384
Reasoning
No
Knowledge cutoff
June 2025
Pricing (standard)
0.45
 AnyTokens in
 / 
1.2
 AnyTokens out

Quickstart

Sample code for Mistral Small 3.2

import requests

url = "https://api.anyapi.ai/v1/chat/completions"

payload = {
    "stream": False,
    "tool_choice": "auto",
    "logprobs": False,
    "model": "Model_Name",
    "messages": [
        {
            "content": [
                {
                    "type": "text",
                    "text": "Hello"
                },
                {
                    "image_url": {
                        "detail": "auto",
                        "url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"
                    },
                    "type": "image_url"
                }
            ],
            "role": "user"
        }
    ]
}
headers = {
    "Authorization": "Bearer AnyAPI_API_KEY",
    "Content-Type": "application/json"
}

response = requests.post(url, json=payload, headers=headers)

print(response.json())
import requests url = "https://api.anyapi.ai/v1/chat/completions" payload = { "stream": False, "tool_choice": "auto", "logprobs": False, "model": "Model_Name", "messages": [ { "content": [ { "type": "text", "text": "Hello" }, { "image_url": { "detail": "auto", "url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg" }, "type": "image_url" } ], "role": "user" } ] } headers = { "Authorization": "Bearer AnyAPI_API_KEY", "Content-Type": "application/json" } response = requests.post(url, json=payload, headers=headers) print(response.json())
View docs
Copy
Code is copied
const url = 'https://api.anyapi.ai/v1/chat/completions';
const options = {
  method: 'POST',
  headers: {Authorization: 'Bearer AnyAPI_API_KEY', 'Content-Type': 'application/json'},
  body: '{"stream":false,"tool_choice":"auto","logprobs":false,"model":"Model_Name","messages":[{"content":[{"type":"text","text":"Hello"},{"image_url":{"detail":"auto","url":"https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"},"type":"image_url"}],"role":"user"}]}'
};

try {
  const response = await fetch(url, options);
  const data = await response.json();
  console.log(data);
} catch (error) {
  console.error(error);
}
const url = 'https://api.anyapi.ai/v1/chat/completions'; const options = { method: 'POST', headers: {Authorization: 'Bearer AnyAPI_API_KEY', 'Content-Type': 'application/json'}, body: '{"stream":false,"tool_choice":"auto","logprobs":false,"model":"Model_Name","messages":[{"content":[{"type":"text","text":"Hello"},{"image_url":{"detail":"auto","url":"https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"},"type":"image_url"}],"role":"user"}]}' }; try { const response = await fetch(url, options); const data = await response.json(); console.log(data); } catch (error) { console.error(error); }
View docs
Copy
Code is copied
curl --request POST \
  --url https://api.anyapi.ai/v1/chat/completions \
  --header 'Authorization: Bearer AnyAPI_API_KEY' \
  --header 'Content-Type: application/json' \
  --data '{
  "stream": false,
  "tool_choice": "auto",
  "logprobs": false,
  "model": "Model_Name",
  "messages": [
    {
      "content": [
        {
          "type": "text",
          "text": "Hello"
        },
        {
          "image_url": {
            "detail": "auto",
            "url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"
          },
          "type": "image_url"
        }
      ],
      "role": "user"
    }
  ]
}'
curl --request POST \ --url https://api.anyapi.ai/v1/chat/completions \ --header 'Authorization: Bearer AnyAPI_API_KEY' \ --header 'Content-Type: application/json' \ --data '{ "stream": false, "tool_choice": "auto", "logprobs": false, "model": "Model_Name", "messages": [ { "content": [ { "type": "text", "text": "Hello" }, { "image_url": { "detail": "auto", "url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg" }, "type": "image_url" } ], "role": "user" } ] }'
View docs
Copy
Code is copied
View docs
Code examples coming soon...

Limitations & Trade-offs

Where Mistral Small 3.2 falls short

1
Not a reasoning model. Mistral Small 3.2 has no dedicated extended-thinking mode and ranks low on agentic and terminal benchmarks such as Tau2-Bench, Terminal-Bench and SciCode. Multi-step planning, long tool-use chains, and hard math or scientific reasoning are where it struggles. For autonomous agents or complex reasoning workloads, a reasoning-tier model is a better fit; use 3.2 for the bounded, well-specified steps within a larger pipeline instead.
2
Text-only output with image input only. The model accepts text and images but returns text exclusively — no image, audio, or video generation. Its vision benchmarks (ChartQA, DocVQA) are strong for document and chart understanding, but some vision metrics regressed marginally versus 3.1. If your workload needs audio, speech, or generated media, this model cannot serve it and you must pair it with dedicated multimodal models.
3
Positioned mid-tier on cost among open-weight peers. Independent analysis notes the model is somewhat more expensive than the median comparable open-weight non-reasoning model, even though it remains inexpensive in absolute terms. For extreme high-volume, cost-critical generation, smaller models like Ministral options may offer better blended pricing and higher throughput. Choose 3.2 when its instruction-following and tool-call reliability justify the modest premium over the cheapest alternatives.
4
Superseded within the lineup. Artificial Analysis lists Mistral Small 3.2 as deprecated in favor of newer Small releases, continuing benchmarking only for a default workload. The weights remain freely usable under Apache 2.0 and the model is production-stable, but teams starting fresh should evaluate whether a newer Small variant offers better reasoning or context for the same role before committing.

Best-Fit Workloads

Where this model earns its place

01

Structured Data Extraction

‍
The model's native JSON-schema support plus improved instruction adherence make it well suited to extracting fields from text and documents into strict formats. Reduced repetition and better format compliance mean fewer malformed responses and less retry logic. This benefits invoice parsing, entity extraction, and classification pipelines where output must validate against a schema on the first pass.

02

Tool-Calling and Function Steps

‍
Mistral upgraded the 3.2 function-calling template specifically for more reliable tool use, particularly via vLLM. This makes it a dependable executor for bounded, well-defined tool-call steps inside larger workflows — API routing, retrieval triggers, and deterministic actions. It is not built for long autonomous agent chains, so keep planning in a reasoning model and delegate individual, specified tool invocations to 3.2.

03

High-Volume Content and Chat

‍
At roughly 152 tokens/second with sub-second time to first token on Mistral's endpoint, and strong conversational scores on Arena Hard v2 and WildBench v2, the model handles high-throughput chat and content generation efficiently. Its reduced infinite-generation rate improves reliability for user-facing assistants. Best where cost-efficiency and responsiveness matter more than frontier reasoning.

04

Document and Chart Understanding

‍
With image input and strong DocVQA (~94.9%) and ChartQA (~87.4%) scores, the model reads documents and charts and returns structured text answers. This fits light document analysis, report summarization from scanned pages, and chart-to-text extraction. Note that output is text only and some vision metrics fluctuated versus 3.1, so validate on your specific document types before scaling.

Pricing in anytokens via AnyAPI
Input
0.45
₳
Output
1.2
₳
Cache write
—
₳
Cache read
—
₳

Integration

Access Mistral Small 3.2 via AnyAPI.ai

Access Mistral Small 3.2 through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate Mistral Small 3.2 without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access Mistral Small 3.2 and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test Mistral Small 3.2 against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use Mistral Small 3.2 from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use Mistral Small 3.2 for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

Mistral Small 3.2 24B is best for structured data extraction, reliable tool calling, high-volume chat, and document or chart understanding. Its 2506 update specifically improves instruction following, reduces repetitive generations, and strengthens function calling, making it dependable for bounded, well-specified production tasks rather than complex autonomous reasoning.

Mistral Small 3.2 24B supports a 128K-token context window, unchanged from Mistral Small 3.1. Hosted providers commonly cap single-request output around 16K tokens. The large context suits long documents and multi-turn sessions, though context window and maximum output are separate limits.

Yes. Mistral Small 3.2 24B accepts both text and image input and returns text only. It supports native function/tool calling with an upgraded template for more reliable tool use, plus structured JSON output via response_format schemas. It does not generate images, audio, or video.

Mistral Small 3.2 24B keeps 3.1's architecture, 128K context, and vision input but improves instruction following (about 82.8% to 84.8% on Mistral's WildBench), roughly halves infinite generations, and upgrades function calling. Conversational scores like Arena Hard v2 more than doubled. It is the safer default for new production deployments.

No. Mistral Small 3.2 24B is a standard instruct model with no dedicated extended-thinking mode, and it ranks low on agentic and terminal benchmarks. For multi-step planning or hard reasoning, use a reasoning-tier model and reserve 3.2 for fast, well-defined steps like extraction and tool calls.

* Benchmark data source: Artificial Analysis artificialanalysis.ai