Google
•
Gemini 3.1 Pro Preview
•
Released 
February 2026

Google
Gemini 3.1 Pro Preview

Google's refined frontier reasoning model with a 1M-token context, tuned for agentic software engineering and grounded factual output.

Modality:
Text
Image
Audio
Video
PDF
model ID
google/gemini-3.1-pro-preview

Output Speed *

124.30
tok/s

Intelligence Index *

29.7
/ 100

Context Window *

1048576
tokens

Input price

12
Anytoken

Output price

72
Anytoken
Gemini 3.1 Pro Preview: Frontier Reasoning With a Million-Token Context and Agentic Tool Use Gemini 3.1 Pro Preview is Google DeepMind's refreshed flagship reasoning model, sitting between the faster Gemini 3 Flash and the research-focused Deep Think tier. It builds on Gemini 3 Pro with better thinking, improved token efficiency, and more grounded, factually consistent output. It accepts text, images, audio, video and PDFs as input, returns text, and carries a 1,048,576-token context window. The model is optimized for software-engineering behavior and agentic workflows that require precise tool usage and reliable multi-step execution, making it well suited to coding agents and long-context analysis. Access Gemini 3.1 Pro Preview via the AnyAPI.ai unified API and start building reasoning and coding agents today.

Performance

Where Gemini 3.1 Pro Preview Earns Its Keep: Reasoning and Agentic Coding

Gemini 3.1 Pro Preview is strongest on abstract reasoning and agentic software engineering. Google reports a verified 77.1% on ARC-AGI-2, more than double Gemini 3 Pro's score, alongside 94.3% on GPQA Diamond and 80.6% on SWE-Bench Verified. These gains matter because the model was explicitly tuned for reliable multi-step tool execution and improved token efficiency, requiring fewer output tokens for comparable results. In production, that translates into more dependable coding agents, fewer wasted reasoning tokens, and more grounded, factually consistent answers—useful for teams already running Gemini 3 Pro who want a drop-in quality upgrade at the same tier.

Benchmarks

Gemini 3.1 Pro Preview Benchmarks: Reasoning Leads, Coding Trails Opus

Independent trackers place Gemini 3.1 Pro Preview near the top of frontier reasoning. On GPQA Diamond it scores roughly 94%, and it reaches the #1 position on LM Arena human-preference rankings. On coding, third-party analysis notes Claude Opus 4.5 still leads SWE-Bench Verified by several points, so Gemini's advantage is clearest on abstract reasoning, science, and multimodal tasks rather than raw autonomous code resolution. Artificial Analysis catalogs strong long-context and coding indices but a more modest agentic index, meaning it is a broad reasoning generalist rather than a pure agent specialist. Scores vary by thinking level and test conditions.

Output Speed

*
124.30
tok/s

Intelligence Index

*
29.7
/ 100

MMLU *

Broad world knowledge and problem-solving
0
%

GPQA *

PhD-level scientific reasoning across physics, biology, chemistry.
94
%

HLE *

Adherence to multi-step structured instructions.
47
%

LiveCodeBench *

Tool-calling reliability in long agentic loops.
0
%

Technical Specifications

What the model supports

Gemini 3.1 Pro Preview accepts text, image, audio, video and PDF input and returns text only. It offers a 1,048,576-token context window with up to 65,536 output tokens per response. The most production-relevant characteristics are the very large context and the thinking_level control: the new medium setting adds a middle ground between speed and reasoning depth, letting teams balance latency and cost per request. The 1M window handles entire codebases or dozens of documents without chunking, but responses remain capped at ~64K tokens, so extremely long single generations still require chaining.
Verified Specifications — 
Gemini 3.1 Pro Preview
*
Input modalities
Text
Image
Audio
Video
PDF
output modalities
Text
Context window
1048576
 tokens
Maximum output tokens
65536
Reasoning
Yes
Knowledge cutoff
February 2026
Pricing (standard)
12
 AnyTokens in
 / 
72
 AnyTokens out

Limitations & Trade-offs

Where Gemini 3.1 Pro Preview falls short

1
Preview status. Gemini 3.1 Pro Preview is a public preview, not a stable GA release, and Google's own model card notes preview evaluations. Preview endpoints can change behavior or be deprecated with limited notice, so building it as a hidden dependency in critical systems calls for careful regression testing and a fallback path before committing production traffic.
2
Coding trails the coding leader. While reasoning and science scores are frontier-level, independent analysis shows Claude Opus 4.5 still leads SWE-Bench Verified by several points. For teams whose primary metric is autonomous resolution of real GitHub issues, a coding-specialized alternative may resolve more issues per run, making Gemini 3.1 Pro a strong generalist rather than the outright best pure coding agent.
3
Output length is capped well below context. The 1M-token context is among the largest available, but responses are limited to 65,536 output tokens. Workloads that need extremely long single generations—full document rewrites, large synthetic datasets, exhaustive code files—still require chaining or streaming across multiple calls, which adds orchestration complexity despite the huge input window.
4
Higher-cost output tier. Gemini 3.1 Pro sits in a premium reasoning tier, and long or high-thinking-level generations accumulate output cost quickly, with rates rising further for prompts above 200,000 tokens. For high-volume, latency-sensitive, or simple tasks, a Flash-class sibling is more economical; reserve Gemini 3.1 Pro for genuinely reasoning-heavy work.

Best-Fit Workloads

Where this model earns its place

01

Agentic software engineering

‍
Gemini 3.1 Pro Preview was explicitly optimized for software-engineering behavior and agentic workflows requiring precise tool usage and reliable multi-step execution. Its dedicated gemini-3.1-pro-preview-customtools endpoint prioritizes custom tools like view_file or search_code, and a strong SWE-Bench Verified score supports real-world coding tasks. It fits IDE agents, automated code review, and multi-step debugging pipelines, though the coding leader may edge it on raw issue-resolution rate.

02

Large-codebase and long-document analysis

‍
The 1,048,576-token context window can absorb entire repositories, lengthy contracts, or dozens of research papers without chunking, and Google reports it comprehends code alongside documentation, PDFs, and mixed media in one prompt. This suits repository-wide refactoring analysis, contract review, and literature synthesis. Remember the ~64K output cap: input scales far beyond what a single response can regenerate.

03

Scientific and technical reasoning
‍

A reported 94.3% on GPQA Diamond—described as the highest recorded—paired with the 1M context makes Gemini 3.1 Pro practical for graduate-level science QA, hypothesis generation, and synthesizing across multiple papers at once. Its more grounded, factually consistent behavior helps technical writing and research assistance where accuracy matters more than raw speed.

04

Autonomous research and browsing agents

‍
Improved agentic reliability and strong tool orchestration, alongside high BrowseComp performance (~85.9%), support multi-step research agents that gather information from multiple sources, verify facts, and produce structured reports with minimal supervision. Combine tool calling with the medium or high thinking level to balance depth against latency; note the agentic index is solid but below dedicated agent specialists.

Pricing in anytokens via AnyAPI
Input
12
₳
Output
72
₳
Cache write
—
₳
Cache read
—
₳

Integration

Access Gemini 3.1 Pro Preview via AnyAPI.ai

Access Gemini 3.1 Pro Preview through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate Gemini 3.1 Pro Preview without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access Gemini 3.1 Pro Preview and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test Gemini 3.1 Pro Preview against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use Gemini 3.1 Pro Preview from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use Gemini 3.1 Pro Preview for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

Gemini 3.1 Pro Preview has a 1,048,576-token context window (about 1M tokens) for input, matching Gemini 3 Pro. It can process entire codebases, long contracts, or dozens of research papers in a single prompt. Maximum output per response is capped at 65,536 tokens, so very long single generations still require chaining.

The Gemini 3.1 Pro Preview API accepts text, images, audio, video and PDF/document input, and returns text output. It does not generate images, audio or video. This native multimodal input lets a single request combine code, documentation, media and long documents within the 1M-token context window.

Gemini 3.1 Pro Preview is a refinement of Gemini 3 Pro with the same 1M context, modalities and pricing tier. It adds better thinking, improved token efficiency, more grounded output, and a new medium thinking level. Benchmarks improve notably—ARC-AGI-2 more than doubles to 77.1%—making it a low-friction quality upgrade for existing users.

Yes. It supports a thinking_level parameter with low, medium and high settings to balance quality, latency and cost; medium is new in 3.1. It also supports function/tool calling and structured outputs, plus a separate gemini-3.1-pro-preview-customtools endpoint that prioritizes custom tools and bash for agentic workflows.

It is strong for coding, scoring around 80.6% on SWE-Bench Verified and 2887 Elo on LiveCodeBench Pro (Google-reported), and it was optimized for agentic software engineering. However, independent analysis notes Claude Opus 4.5 still leads SWE-Bench Verified by several points, so it is an excellent generalist coder rather than the single best pure coding agent.

* Benchmark data source: Artificial Analysis artificialanalysis.ai