Mistral AI
•
codestral-2508
•
Released 
August 2025

Mistral AI
codestral-2508

Mistral's coding specialist tuned for low-latency fill-in-the-middle completion and high-frequency IDE autocomplete across a 256K context.

Modality:
Text
PDF
model ID
mistralai/codestral-2508

Output Speed *

0.00
tok/s

Intelligence Index *

0
/ 100

Context Window *

256000
tokens

Input price

1.8
Anytoken

Output price

5.4
Anytoken
Codestral 2508: Low-Latency Code Completion Built for the IDE Loop Codestral 2508 (Codestral 25.08) is Mistral's code-specialized language model, released in August 2025 as the completion engine of Mistral's coding stack alongside Devstral for agents and Codestral Embed for retrieval. It is optimized for low-latency, high-frequency tasks: fill-in-the-middle (FIM), code correction, and test generation. Trained across 80+ programming languages and paired with a 256K-token context window, it targets inline autocomplete and refactoring where suggestion speed and codebase awareness matter more than deep chain-of-thought reasoning. It is the strongest fit for editor-integrated completion rather than autonomous multi-step engineering. Integrate Codestral 2508 through the AnyAPI.ai API

Performance

Where Codestral 2508 Earns Its Place: Autocomplete and FIM

Codestral 2508 is built for the tight edit-suggest loop inside an IDE rather than for open-ended reasoning. Mistral positions it for low-latency fill-in-the-middle, code correction, and test generation, and the Codestral line reports higher completion accept rates and improved suggestion reliability versus earlier versions. In practice this means fast, context-aware inline completions that developers accept without breaking flow. The 256K context lets it read surrounding files instead of a single buffer, improving suggestion relevance. The consequence: it excels as a completion backend, but it is not designed to plan or execute autonomous multi-step tasks.

Benchmarks

Independent Benchmark Signals for Codestral 2508

Independent, model-specific benchmark coverage for Codestral 2508 is thin. Several aggregators note that no verified third-party coding-accuracy scores are consistently published for this exact version, so its resolve rate on suites like SWE-bench Verified cannot be stated with confidence. Where scores are attributed, they are inconsistent across sources and reasoning settings, so they are not reliable for selection. The practical takeaway: evaluate Codestral 2508 on your own repositories for completion acceptance, latency, and correctness rather than relying on published leaderboard figures, which currently do not cleanly isolate this release.

Output Speed

*
0.00
tok/s

Intelligence Index

*
0
/ 100

MMLU *

Broad world knowledge and problem-solving
0
%

GPQA *

PhD-level scientific reasoning across physics, biology, chemistry.
0
%

HLE *

Adherence to multi-step structured instructions.
0
%

LiveCodeBench *

Tool-calling reliability in long agentic loops.
0
%

Technical Specifications

What the model supports

Codestral 2508 is a text-and-document input, text-output coding model with a 256K-token context window and a comparatively short maximum output, commonly reported around 4K tokens by hosting providers. It exposes chat completions plus a dedicated FIM endpoint, function calling, structured outputs, prompt caching, and batching. The most production-relevant characteristics are the large context, which lets it read multi-file surroundings, and the modest output ceiling, which favors completions and patches over long generated files. It does not provide a dedicated extended-reasoning mode.
Verified Specifications — 
codestral-2508
*
Input modalities
Text
PDF
output modalities
Text
Context window
256000
 tokens
Maximum output tokens
204800
Reasoning
No
Knowledge cutoff
August 2025
Pricing (standard)
1.8
 AnyTokens in
 / 
5.4
 AnyTokens out

Quickstart

Sample code for codestral-2508

import requests

url = "https://api.anyapi.ai/v1/chat/completions"

payload = {
    "stream": False,
    "tool_choice": "auto",
    "logprobs": False,
    "model": "codestral-2508",
    "messages": [
        {
            "role": "user",
            "content": "Hello"
        }
    ]
}
headers = {
    "Authorization": "Bearer AnyAPI_API_KEY",
    "Content-Type": "application/json"
}

response = requests.post(url, json=payload, headers=headers)

print(response.json())
import requests url = "https://api.anyapi.ai/v1/chat/completions" payload = { "stream": False, "tool_choice": "auto", "logprobs": False, "model": "codestral-2508", "messages": [ { "role": "user", "content": "Hello" } ] } headers = { "Authorization": "Bearer AnyAPI_API_KEY", "Content-Type": "application/json" } response = requests.post(url, json=payload, headers=headers) print(response.json())
View docs
Copy
Code is copied
const url = 'https://api.anyapi.ai/v1/chat/completions';
const options = {
  method: 'POST',
  headers: {Authorization: 'Bearer AnyAPI_API_KEY', 'Content-Type': 'application/json'},
  body: '{"stream":false,"tool_choice":"auto","logprobs":false,"model":"codestral-2508","messages":[{"role":"user","content":"Hello"}]}'
};

try {
  const response = await fetch(url, options);
  const data = await response.json();
  console.log(data);
} catch (error) {
  console.error(error);
}
const url = 'https://api.anyapi.ai/v1/chat/completions'; const options = { method: 'POST', headers: {Authorization: 'Bearer AnyAPI_API_KEY', 'Content-Type': 'application/json'}, body: '{"stream":false,"tool_choice":"auto","logprobs":false,"model":"codestral-2508","messages":[{"role":"user","content":"Hello"}]}' }; try { const response = await fetch(url, options); const data = await response.json(); console.log(data); } catch (error) { console.error(error); }
View docs
Copy
Code is copied
curl --request POST \
  --url https://api.anyapi.ai/v1/chat/completions \
  --header 'Authorization: Bearer AnyAPI_API_KEY' \
  --header 'Content-Type: application/json' \
  --data '{
  "stream": false,
  "tool_choice": "auto",
  "logprobs": false,
  "model": "codestral-2508",
  "messages": [
    {
      "role": "user",
      "content": "Hello"
    }
  ]
}'
curl --request POST \ --url https://api.anyapi.ai/v1/chat/completions \ --header 'Authorization: Bearer AnyAPI_API_KEY' \ --header 'Content-Type: application/json' \ --data '{ "stream": false, "tool_choice": "auto", "logprobs": false, "model": "codestral-2508", "messages": [ { "role": "user", "content": "Hello" } ] }'
View docs
Copy
Code is copied
View docs
Code examples coming soon...

Limitations & Trade-offs

Where codestral-2508 falls short

1
Short maximum output. Multiple hosting providers report an output ceiling around 4K tokens, which is well suited to completions and patches but constraining for generating large files, extensive multi-function modules, or long documentation in one pass. Workloads that need long single-shot generation will require chunking or a model with a higher output limit; for those, a general-purpose or agentic model is preferable.
2
No dedicated reasoning mode. Codestral 2508 is tuned for low-latency completion and does not expose an extended chain-of-thought or reasoning-effort control. For complex algorithmic problem solving, architectural planning, or tasks that benefit from deliberate step-by-step reasoning, a reasoning-capable model or Mistral's agentic Devstral line is the better choice.
3
Sparse independent benchmarks for this exact version. Several aggregators note that verified third-party coding-accuracy scores for Codestral 2508 specifically are not consistently published, and figures that do appear vary by source and test conditions. Teams cannot rely on leaderboard numbers to predict behavior and should validate acceptance rate and correctness against their own codebases before wider rollout.
4
Not an autonomous agent. The model generates and completes code but is not designed to plan and execute multi-step engineering workflows across files on its own — that role belongs to Devstral within Mistral's stack. Using Codestral 2508 alone for autonomous issue resolution or cross-file refactoring will underperform a purpose-built agent.

Best-Fit Workloads

Where this model earns its place

01

IDE Autocomplete Backend

‍
Codestral 2508's core design target: low-latency inline completion inside editors. Mistral positions it for exactly this loop, and the Codestral line reports improved completion accept rates and suggestion reliability. Its 256K context lets it read surrounding files for relevant suggestions. Production copilots and plugin integrations benefit most; the short output ceiling is a non-issue because completions are inherently small.

02

Fill-in-the-Middle Editing

‍
FIM lets the model generate code between existing lines rather than only appending, and Codestral 2508 ships a dedicated FIM endpoint. This suits refactoring inside functions, inserting missing logic, and patch-style edits where surrounding context constrains the output. It is a defining capability of this model rather than a general LLM feature, making it a strong fit for structured in-place code edits.

03

Code Correction and Test Generation

‍
Mistral highlights code correction and automated test generation as primary use cases. The model's low-latency profile and language breadth suit pre-commit fix suggestions, linting-adjacent repair, and scaffolding unit tests in CI. Outputs are small and targeted, aligning with the short output ceiling. Validate generated tests for correctness, since coverage and edge-case handling still require review.

04

Large-Context Code Understanding

‍
The 256K-token context window lets Codestral 2508 read across large files or multiple related files in a single request, improving suggestion relevance beyond single-buffer models. This benefits monorepo-adjacent completion and context-aware refactoring. Note that provider documentation surfaces differing context figures for the endpoint, so confirm the effective context limit with your chosen host before relying on the full window.

Pricing in anytokens via AnyAPI
Input
1.8
₳
Output
5.4
₳
Cache write
—
₳
Cache read
—
₳

Integration

Access codestral-2508 via AnyAPI.ai

Access codestral-2508 through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate codestral-2508 without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access codestral-2508 and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test codestral-2508 against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use codestral-2508 from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use codestral-2508 for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

Codestral 2508 is Mistral's coding model optimized for low-latency, high-frequency tasks: fill-in-the-middle completion, code correction, and test generation. It is designed as an IDE autocomplete backend rather than an autonomous coding agent, making it best for fast inline suggestions and in-place code edits across 80+ programming languages.

Codestral 2508 is widely reported with a 256,000-token context window, allowing it to read large files or multiple related files at once for more relevant completions. Note that some provider documentation surfaces a different figure for the API endpoint, so confirm the effective context limit with your chosen host.

Yes. Codestral 2508 accepts tools and tool_choice for function calling and supports structured outputs via a JSON schema in response_format. It also exposes a dedicated fill-in-the-middle endpoint plus prefix, predicted outputs, document QnA, and batching, and uses an OpenAI-compatible request format.

No. Codestral 2508 is tuned for low-latency completion and does not expose a dedicated extended-reasoning or reasoning-effort control. For complex algorithmic reasoning, architectural planning, or autonomous multi-step engineering, a reasoning-capable model or Mistral's agentic Devstral line is a better fit.

Codestral 2508 is the completion and fill-in-the-middle engine for fast inline suggestions in the editor. Devstral is the agentic model built to scan codebases, edit across files, and draft pull requests. They are complementary: use Codestral for the keystroke loop and Devstral for delegated multi-step engineering tasks.

* Benchmark data source: Artificial Analysis artificialanalysis.ai