Google's refined frontier reasoning model with a 1M-token context, tuned for agentic software engineering and grounded factual output.
Output Speed *
Intelligence Index *
Context Window *
Input price
Output price
Performance
Where Gemini 3.1 Pro Preview Earns Its Keep: Reasoning and Agentic Coding
Benchmarks
Gemini 3.1 Pro Preview Benchmarks: Reasoning Leads, Coding Trails Opus
Output Speed
Intelligence Index
MMLU *
GPQA *
HLE *
LiveCodeBench *
Technical Specifications
What the model supports
Limitations & Trade-offs
Best-Fit Workloads
Where this model earns its place
Agentic software engineering
Gemini 3.1 Pro Preview was explicitly optimized for software-engineering behavior and agentic workflows requiring precise tool usage and reliable multi-step execution. Its dedicated gemini-3.1-pro-preview-customtools endpoint prioritizes custom tools like view_file or search_code, and a strong SWE-Bench Verified score supports real-world coding tasks. It fits IDE agents, automated code review, and multi-step debugging pipelines, though the coding leader may edge it on raw issue-resolution rate.
Large-codebase and long-document analysis
The 1,048,576-token context window can absorb entire repositories, lengthy contracts, or dozens of research papers without chunking, and Google reports it comprehends code alongside documentation, PDFs, and mixed media in one prompt. This suits repository-wide refactoring analysis, contract review, and literature synthesis. Remember the ~64K output cap: input scales far beyond what a single response can regenerate.
Scientific and technical reasoning
A reported 94.3% on GPQA Diamond—described as the highest recorded—paired with the 1M context makes Gemini 3.1 Pro practical for graduate-level science QA, hypothesis generation, and synthesizing across multiple papers at once. Its more grounded, factually consistent behavior helps technical writing and research assistance where accuracy matters more than raw speed.
Autonomous research and browsing agents
Improved agentic reliability and strong tool orchestration, alongside high BrowseComp performance (~85.9%), support multi-step research agents that gather information from multiple sources, verify facts, and produce structured reports with minimal supervision. Combine tool calling with the medium or high thinking level to balance depth against latency; note the agentic index is solid but below dedicated agent specialists.