Google
•
Gemini 2.5 Pro
•
Released 
June 2025

Google
Gemini 2.5 Pro

Google's flagship reasoning model with a 1M-token context window for whole-codebase and long-document analysis.

Modality:
Text
Image
Audio
Video
PDF
model ID
google/gemini-2.5-pro

Output Speed *

0.00
tok/s

Intelligence Index *

16.1
/ 100

Context Window *

1048576
tokens

Input price

13.5
Anytoken

Output price

90
Anytoken
Gemini 2.5 Pro: Million-Token Reasoning for Codebases and Documents Gemini 2.5 Pro is Google DeepMind's flagship reasoning model and the top tier of the Gemini 2.5 family, above Gemini 2.5 Flash and Flash-Lite. It is a natively multimodal "thinking" model that reasons through problems before answering. Its defining trait is a 1,000,000-token context window paired with strong reasoning, which lets it analyze entire codebases, large document sets, and mixed text, image, audio, and video inputs in a single request. It fits workloads where deep reasoning over very long inputs matters more than raw output speed or lowest cost. Start building with the Gemini 2.5 Pro API on AnyAPI.ai

Performance

Reasoning Depth Across Very Long Inputs

Gemini 2.5 Pro's strength is sustained reasoning over large, mixed-modality inputs rather than short-turn speed. It integrates a native "thinking" step before responding, and Google's technical report shows performance on AIME 2025 of 88.0% and GPQA diamond of 86.4% for Gemini 2.5 Pro, a large jump over the 1.5 generation. For developers this means the model can reason across an entire codebase or document corpus in one pass, reducing chunking and RAG plumbing. The trade-off: thinking plus long context increases latency and output cost, so it suits deliberate analysis over high-throughput serving.

Benchmarks

How Gemini 2.5 Pro Scores on Independent Evaluations

On independent testing, Gemini 2.5 Pro is strong in science and math reasoning but not the clear coding leader. It scores <cite index="20-3">84.0% on GPQA Diamond</cite> and <cite index="20-4">92.0% on AIME 2024 for single attempt</cite>. On coding, <cite index="22-1">Claude 3.7 Sonnet maintains an edge in SWE-bench Verified (70.3% vs. 63.8%), and o3-mini leads slightly in LiveCodeBench v5 (74.1% vs. 70.4%)</cite>. On throughput, Artificial Analysis reports <cite index="17-1">143 tokens per second</cite>, though other trackers report lower figures depending on configuration. Treat these as independent measurements, not official specs.

Output Speed

*
0.00
tok/s

Intelligence Index

*
16.1
/ 100

MMLU *

Broad world knowledge and problem-solving
86
%

GPQA *

PhD-level scientific reasoning across physics, biology, chemistry.
84
%

HLE *

Adherence to multi-step structured instructions.
23
%

LiveCodeBench *

Tool-calling reliability in long agentic loops.
80
%

Technical Specifications

What the model supports

Gemini 2.5 Pro is a proprietary reasoning model with a ~1M-token context window and up to ~64K output tokens per response. It accepts text, images, PDFs/documents, audio, and video, and returns text only — it is not an image or audio generator. The two specs with the greatest production impact are the large context window, which removes most chunking work, and the 64K output cap, which is generous for full documents or code files but still bounds single-call generation. Reasoning is native, with configurable thinking budgets exposed via the API.
Verified Specifications — 
Gemini 2.5 Pro
*
Input modalities
Text
Image
Audio
Video
PDF
output modalities
Text
Context window
1048576
 tokens
Maximum output tokens
65536
Reasoning
Yes
Knowledge cutoff
June 2025
Pricing (standard)
13.5
 AnyTokens in
 / 
90
 AnyTokens out

Quickstart

Sample code for Gemini 2.5 Pro

import requests

url = "https://api.anyapi.ai/v1/chat/completions"

payload = {
    "stream": False,
    "tool_choice": "auto",
    "logprobs": False,
    "model": "gemini-2.5-pro",
    "messages": [
        {
            "content": [
                {
                    "type": "text",
                    "text": "Hello"
                },
                {
                    "image_url": {
                        "detail": "auto",
                        "url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"
                    },
                    "type": "image_url"
                }
            ],
            "role": "user"
        }
    ]
}
headers = {
    "Authorization": "Bearer AnyAPI_API_KEY",
    "Content-Type": "application/json"
}

response = requests.post(url, json=payload, headers=headers)

print(response.json())
import requests url = "https://api.anyapi.ai/v1/chat/completions" payload = { "stream": False, "tool_choice": "auto", "logprobs": False, "model": "gemini-2.5-pro", "messages": [ { "content": [ { "type": "text", "text": "Hello" }, { "image_url": { "detail": "auto", "url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg" }, "type": "image_url" } ], "role": "user" } ] } headers = { "Authorization": "Bearer AnyAPI_API_KEY", "Content-Type": "application/json" } response = requests.post(url, json=payload, headers=headers) print(response.json())
View docs
Copy
Code is copied
const url = 'https://api.anyapi.ai/v1/chat/completions';
const options = {
  method: 'POST',
  headers: {Authorization: 'Bearer AnyAPI_API_KEY', 'Content-Type': 'application/json'},
  body: '{"stream":false,"tool_choice":"auto","logprobs":false,"model":"gemini-2.5-pro","messages":[{"content":[{"type":"text","text":"Hello"},{"image_url":{"detail":"auto","url":"https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"},"type":"image_url"}],"role":"user"}]}'
};

try {
  const response = await fetch(url, options);
  const data = await response.json();
  console.log(data);
} catch (error) {
  console.error(error);
}
const url = 'https://api.anyapi.ai/v1/chat/completions'; const options = { method: 'POST', headers: {Authorization: 'Bearer AnyAPI_API_KEY', 'Content-Type': 'application/json'}, body: '{"stream":false,"tool_choice":"auto","logprobs":false,"model":"gemini-2.5-pro","messages":[{"content":[{"type":"text","text":"Hello"},{"image_url":{"detail":"auto","url":"https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"},"type":"image_url"}],"role":"user"}]}' }; try { const response = await fetch(url, options); const data = await response.json(); console.log(data); } catch (error) { console.error(error); }
View docs
Copy
Code is copied
curl --request POST \
  --url https://api.anyapi.ai/v1/chat/completions \
  --header 'Authorization: Bearer AnyAPI_API_KEY' \
  --header 'Content-Type: application/json' \
  --data '{
  "stream": false,
  "tool_choice": "auto",
  "logprobs": false,
  "model": "gemini-2.5-pro",
  "messages": [
    {
      "content": [
        {
          "type": "text",
          "text": "Hello"
        },
        {
          "image_url": {
            "detail": "auto",
            "url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"
          },
          "type": "image_url"
        }
      ],
      "role": "user"
    }
  ]
}'
curl --request POST \ --url https://api.anyapi.ai/v1/chat/completions \ --header 'Authorization: Bearer AnyAPI_API_KEY' \ --header 'Content-Type: application/json' \ --data '{ "stream": false, "tool_choice": "auto", "logprobs": false, "model": "gemini-2.5-pro", "messages": [ { "content": [ { "type": "text", "text": "Hello" }, { "image_url": { "detail": "auto", "url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg" }, "type": "image_url" } ], "role": "user" } ] }'
View docs
Copy
Code is copied
View docs
Code examples coming soon...

Limitations & Trade-offs

Where Gemini 2.5 Pro falls short

1
Latency under long context and thinking. Because Gemini 2.5 Pro runs a native reasoning step and can ingest very large prompts, response times grow with input size and thinking budget. Independent reports note that very long prompts produce noticeable time-to-first-token delays. This matters for interactive chat or real-time UIs; for those, a faster model like Gemini 2.5 Flash is usually the better fit, reserving Pro for deliberate analysis.
2
Output cost and length ceiling. The model tops out around 64K output tokens per response and sits in a higher-cost tier for generation than the Flash siblings. Long, verbose outputs — full documentation sets or large code files — are comparatively expensive and may need multiple calls to exceed the cap. For high-volume generation, a cheaper family member or another provider is more economical.
3
Not the strongest coding leader. On agentic coding, independent benchmarks place Gemini 2.5 Pro behind Claude 3.7 Sonnet on SWE-bench Verified (63.8% vs 70.3%) and slightly behind o3-mini on LiveCodeBench v5. It remains a capable coder with a large context advantage, but for pure code-editing agents where SWE-bench performance drives selection, a competitor may edge it out.
4
Text-only output and 2M context caveats. Gemini 2.5 Pro accepts rich multimodal input but returns only text — it cannot generate images or audio. The headline 2M-token context is available on select Vertex AI enterprise tiers, not the standard 1M API window, so plan for 1M unless you have that tier. Long-context recall also varies by task, so validate retrieval on your own corpus before relying on it.

Best-Fit Workloads

Where this model earns its place

01

Whole-codebase analysis and refactoring

‍
The 1M-token window lets Gemini 2.5 Pro read an entire repository — reportedly tens of thousands of lines — in a single prompt, then reason about architecture, cross-file dependencies, and refactors without chunking or RAG. This is the model's most characteristic strength. It suits code review, migration planning, and large-scale debugging where seeing the whole project at once matters. For narrow, high-precision code-editing agents, benchmark leaders like Claude 3.7 Sonnet may perform slightly better.

02

Long-document and multi-source research

‍
Combining reasoning with a million-token context makes Gemini 2.5 Pro well suited to synthesizing lengthy contracts, filings, research papers, or entire documentation sets in one pass. It can hold multiple sources in context and draw cross-document conclusions, reducing the retrieval plumbing a smaller-context model requires. Validate long-context recall on your own corpus, since practical retrieval at extreme lengths varies by task.

03

Multimodal document and media understanding

‍
Native input support for text, images, PDFs, audio, and video lets the model handle mixed-media inputs — diagrams alongside code, or transcripts alongside slides — in a single request. It reports strong image-understanding scores (e.g. MMMU 81.7% in independent testing). Useful for technical document extraction, screen and diagram interpretation, and media analysis. Note that output is text only; it cannot produce images or audio.

04

Complex reasoning, math, and science

‍
With configurable thinking and high scores on GPQA Diamond and AIME, Gemini 2.5 Pro fits STEM-heavy workloads: scientific analysis, structured mathematical reasoning, and multi-step problem solving. The thinking budget lets you trade latency and cost for accuracy on hard problems. For simple, high-volume queries the reasoning overhead is unnecessary — use a lighter model there.

Pricing in anytokens via AnyAPI
Input
13.5
₳
Output
90
₳
Cache write
2.25
₳
Cache read
0.75
₳

Integration

Access Gemini 2.5 Pro via AnyAPI.ai

Access Gemini 2.5 Pro through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate Gemini 2.5 Pro without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access Gemini 2.5 Pro and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test Gemini 2.5 Pro against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use Gemini 2.5 Pro from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use Gemini 2.5 Pro for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

Gemini 2.5 Pro has a context window of 1,048,576 tokens (about 1M) on the standard API, with up to about 64K output tokens per response. Select Vertex AI enterprise tiers extend the context to 2 million tokens. The large window lets it process entire codebases or large document sets without chunking.

Yes, but with nuance. It is a strong coder — especially valuable for whole-codebase analysis thanks to its 1M-token context. On independent agentic-coding benchmarks it scores 63.8% on SWE-bench Verified, behind Claude 3.7 Sonnet (70.3%). For large-context code understanding it excels; for pure code-editing agents a competitor may edge it out.

Yes. Gemini 2.5 Pro is a native "thinking" model with configurable thinking budgets exposed through the API's thinking_config parameter, letting you cap internal reasoning tokens. It can also return inspectable thought summaries with headers, key details, and tool-call annotations to help you audit how it reached an answer.

Gemini 2.5 Pro accepts text, images, PDFs and other documents, audio, and video as input within a single request. It returns text only — it does not generate images or audio. This makes it suited to multimodal understanding and extraction rather than media generation.

Gemini 2.5 Pro is generally available through the Gemini API, Google AI Studio, and Vertex AI, with SDKs for Python, JavaScript, Go, and REST. It also supports function calling and native MCP. Through AnyAPI.ai you can access it via a unified API alongside other models without a separate Google integration.

* Benchmark data source: Artificial Analysis artificialanalysis.ai