Google
•
Gemini 2.5 Flash (Non-reasoning)
•
Released 
June 2025

Google
Gemini 2.5 Flash (Non-reasoning)

Google's high-throughput workhorse with configurable thinking, 1M-token context, and multimodal input for cost-sensitive production apps.

Modality:
Text
Image
Audio
Video
PDF
model ID
google/gemini-2.5-flash

Output Speed *

0.00
tok/s

Intelligence Index *

9.9
/ 100

Context Window *

1048576
tokens

Input price

1.8
Anytoken

Output price

15
Anytoken
Gemini 2.5 Flash: High-Speed Multimodal Inference with Configurable Thinking Gemini 2.5 Flash is Google's price-performance workhorse in the Gemini 2.5 family, sitting below Gemini 2.5 Pro and above Flash-Lite. It was Google's first Flash model with configurable thinking, letting you trade reasoning depth for latency and cost per request. It accepts text, images, audio, video, and documents, returns text, and handles a 1M-token context window. Its defining trait is throughput: fast token generation and low time-to-first-token make it a strong default for high-volume chat, extraction, and multimodal pipelines that need reasonable intelligence without premium-tier cost. Start building with Gemini 2.5 Flash through the AnyAPI.ai unified API.

Performance

Throughput-First Inference with Optional Reasoning Depth

Gemini 2.5 Flash is built for speed at scale. Independent testing measures output around 223 tokens per second with sub-second time-to-first-token on Google AI Studio, ranking it among the fastest production models. Configurable thinking lets you enable extended reasoning only when a task needs it, and the September 2025 update raised SWE-Bench Verified to 54% while cutting tokens used per task. For developers, this means one endpoint can serve both latency-sensitive chat and heavier agentic steps, keeping infrastructure simple while controlling cost by capping thinking budget per request.

Benchmarks

Where Gemini 2.5 Flash Lands on Independent Benchmarks

On the Artificial Analysis Intelligence Index, Gemini 2.5 Flash in reasoning mode scores 27, placing it above average for its class, while remaining notably fast at roughly 223 tokens per second. Vals AI ranks it #3 of 38 on MMMU Pro and #6 of 20 on SWE-bench Verified for the thinking model, noting competitive performance at a fraction of larger foundation models' cost. One caveat: independent testers flag it as verbose, generating well above the median output volume on the Intelligence Index, which can inflate end-to-end latency and per-task cost even though per-token speed is high.

Output Speed

*
0.00
tok/s

Intelligence Index

*
9.9
/ 100

MMLU *

Broad world knowledge and problem-solving
81
%

GPQA *

PhD-level scientific reasoning across physics, biology, chemistry.
68
%

HLE *

Adherence to multi-step structured instructions.
5
%

LiveCodeBench *

Tool-calling reliability in long agentic loops.
50
%

Technical Specifications

What the model supports

Gemini 2.5 Flash accepts text, code, images, audio, and video and returns text only. It offers a 1,048,576-token context window with up to 65,535 output tokens per response. The most production-relevant feature is configurable thinking: reasoning can be enabled, budgeted, or turned off per request, directly controlling latency and cost. It supports function calling, structured JSON outputs, streaming, context caching, and batch processing. The large context suits long-document and multi-file workloads, but remember input and output share the token budget, so very long inputs reduce room for generation.
Verified Specifications — 
Gemini 2.5 Flash (Non-reasoning)
*
Input modalities
Text
Image
Audio
Video
PDF
output modalities
Text
Context window
1048576
 tokens
Maximum output tokens
65535
Reasoning
Yes
Knowledge cutoff
June 2025
Pricing (standard)
1.8
 AnyTokens in
 / 
15
 AnyTokens out

Quickstart

Sample code for Gemini 2.5 Flash (Non-reasoning)

import requests

url = "https://api.anyapi.ai/v1/chat/completions"

payload = {
    "stream": False,
    "tool_choice": "auto",
    "logprobs": False,
    "model": "gemini-2.5-flash",
    "messages": [
        {
            "content": [
                {
                    "type": "text",
                    "text": "Hello"
                },
                {
                    "image_url": {
                        "detail": "auto",
                        "url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"
                    },
                    "type": "image_url"
                }
            ],
            "role": "user"
        }
    ]
}
headers = {
    "Authorization": "Bearer AnyAPI_API_KEY",
    "Content-Type": "application/json"
}

response = requests.post(url, json=payload, headers=headers)

print(response.json())
import requests url = "https://api.anyapi.ai/v1/chat/completions" payload = { "stream": False, "tool_choice": "auto", "logprobs": False, "model": "gemini-2.5-flash", "messages": [ { "content": [ { "type": "text", "text": "Hello" }, { "image_url": { "detail": "auto", "url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg" }, "type": "image_url" } ], "role": "user" } ] } headers = { "Authorization": "Bearer AnyAPI_API_KEY", "Content-Type": "application/json" } response = requests.post(url, json=payload, headers=headers) print(response.json())
View docs
Copy
Code is copied
const url = 'https://api.anyapi.ai/v1/chat/completions';
const options = {
  method: 'POST',
  headers: {Authorization: 'Bearer AnyAPI_API_KEY', 'Content-Type': 'application/json'},
  body: '{"stream":false,"tool_choice":"auto","logprobs":false,"model":"gemini-2.5-flash","messages":[{"content":[{"type":"text","text":"Hello"},{"image_url":{"detail":"auto","url":"https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"},"type":"image_url"}],"role":"user"}]}'
};

try {
  const response = await fetch(url, options);
  const data = await response.json();
  console.log(data);
} catch (error) {
  console.error(error);
}
const url = 'https://api.anyapi.ai/v1/chat/completions'; const options = { method: 'POST', headers: {Authorization: 'Bearer AnyAPI_API_KEY', 'Content-Type': 'application/json'}, body: '{"stream":false,"tool_choice":"auto","logprobs":false,"model":"gemini-2.5-flash","messages":[{"content":[{"type":"text","text":"Hello"},{"image_url":{"detail":"auto","url":"https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"},"type":"image_url"}],"role":"user"}]}' }; try { const response = await fetch(url, options); const data = await response.json(); console.log(data); } catch (error) { console.error(error); }
View docs
Copy
Code is copied
curl --request POST \
  --url https://api.anyapi.ai/v1/chat/completions \
  --header 'Authorization: Bearer AnyAPI_API_KEY' \
  --header 'Content-Type: application/json' \
  --data '{
  "stream": false,
  "tool_choice": "auto",
  "logprobs": false,
  "model": "gemini-2.5-flash",
  "messages": [
    {
      "content": [
        {
          "type": "text",
          "text": "Hello"
        },
        {
          "image_url": {
            "detail": "auto",
            "url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"
          },
          "type": "image_url"
        }
      ],
      "role": "user"
    }
  ]
}'
curl --request POST \ --url https://api.anyapi.ai/v1/chat/completions \ --header 'Authorization: Bearer AnyAPI_API_KEY' \ --header 'Content-Type: application/json' \ --data '{ "stream": false, "tool_choice": "auto", "logprobs": false, "model": "gemini-2.5-flash", "messages": [ { "content": [ { "type": "text", "text": "Hello" }, { "image_url": { "detail": "auto", "url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg" }, "type": "image_url" } ], "role": "user" } ] }'
View docs
Copy
Code is copied
View docs
Code examples coming soon...

Comparison

Gemini 2.5 Flash vs Gemini 2.5 Pro: Which Tier Fits Your Workload?

Gemini 2.5 Flash and Gemini 2.5 Pro share the same family, 1M-token context, multimodal input, thinking capability, and tool-calling API, so migration between them is straightforward. The practical decision is intelligence versus throughput and cost. Pro targets the hardest reasoning, coding, and long-context analysis, but generates far slower and at a materially higher cost tier. Flash generates several times faster with much lower latency and occupies a lower-cost tier, while still handling reasoning when its thinking budget is enabled. Most teams route routine traffic to Flash and escalate only difficult cases to Pro.

Dimension
Gemini 2.5 Flash (Non-reasoning)
Gemini 2.5 Pro
Context window *
1048576
tokens
1048576
tokens
Output speed *
0.00
tok/s
0.00
tok/s
Intelligence Index *
9.9
16.1
Input pricing
1.8
AnyToken
13.5
AnyToken
Output pricing
15
AnyToken
90
AnyToken
Knowledge cutoff *
June 2025
June 2025

Choose Gemini 2.5 Flash when throughput, low latency, and cost efficiency matter most: high-volume chat, extraction, multimodal pipelines, and agentic loops where per-request speed drives user experience. Choose Gemini 2.5 Pro when a task demands maximum reasoning quality, the most reliable long-context analysis, or the strongest coding accuracy, and you can absorb slower generation and higher cost. A common production pattern uses Flash as the default and escalates only the hardest requests to Pro, balancing quality against spend.

Limitations & Trade-offs

Where Gemini 2.5 Flash (Non-reasoning) falls short

1
Verbose output inflates cost and latency. Independent testing found Gemini 2.5 Flash generates well above the median output volume on the Intelligence Index, nearly double the average for equivalent tasks. Even though per-token speed is high, more tokens per task means higher end-to-end latency and cost. For budget-sensitive or tightly latency-bound applications, constrain thinking budget and set explicit output limits, or consider a more concise model for short-answer workloads.
2
Text-only output. Gemini 2.5 Flash accepts rich multimodal input (images, audio, video, documents) but returns text only. It cannot generate images, audio, or video. If your pipeline needs generated media, pair it with a dedicated generation model; Flash's role is understanding and text generation, not media synthesis.
3
Not the top of the family for hardest reasoning. Flash sits below Gemini 2.5 Pro. On the Artificial Analysis Intelligence Index its reasoning-mode score of 27 is above average but below Pro-tier models. For the most demanding scientific, mathematical, or long-context analytical work where accuracy is critical, Pro or another frontier model remains the safer choice.
4
Shared input/output token budget. The 1M-token context and 65,535-token maximum output draw from the same window. Very long inputs reduce room for generation, so document-heavy prompts that also need long responses must be budgeted carefully. For extreme long-context work, context caching helps control repeat-input cost but does not remove the shared-budget constraint.

Best-Fit Workloads

Where this model earns its place

01

High-volume multimodal extraction

‍
Flash ingests PDFs, images, audio, and video and returns structured text, making it a strong fit for document parsing, receipt and invoice extraction, and media classification at scale. Its fast token generation and low time-to-first-token keep throughput high across large batches, and JSON schema support produces machine-parseable output. Constrain thinking budget for simple extraction to minimize cost per document.

02

Latency-sensitive chat and assistants

‍
With sub-second time-to-first-token and roughly 223 tokens per second, Flash delivers responsive streaming for user-facing chat. Disable or minimize thinking for conversational turns to keep responses immediate, and reserve larger thinking budgets for occasional complex queries. The 1M-token context supports long conversation history and document grounding within a single session.

03

Agentic tool-use loops

‍
The September 2025 update improved agentic tool use, raising SWE-Bench Verified to 54% while using fewer tokens. Combined with function calling and configurable thinking, Flash suits multi-step agents that call tools, evaluate results, and iterate. Its speed keeps agent loops fast, and thinking can be raised on hard planning steps. For the most difficult autonomous coding tasks, escalate to a Pro-tier model.

04

Long-document and RAG grounding

‍
The 1M-token context lets Flash read large documents, multi-file inputs, or extended histories in a single request, reducing chunking complexity. With context caching, repeatedly querying the same large corpus becomes more economical. It fits chat-with-your-data and retrieval-augmented pipelines that need fast responses over large grounding material, though very long inputs reduce available output tokens.

Pricing in anytokens via AnyAPI
Input
1.8
₳
Output
15
₳
Cache write
0.5
₳
Cache read
0.18
₳

Integration

Access Gemini 2.5 Flash (Non-reasoning) via AnyAPI.ai

Access Gemini 2.5 Flash (Non-reasoning) through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate Gemini 2.5 Flash (Non-reasoning) without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access Gemini 2.5 Flash (Non-reasoning) and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test Gemini 2.5 Flash (Non-reasoning) against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use Gemini 2.5 Flash (Non-reasoning) from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use Gemini 2.5 Flash (Non-reasoning) for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

Gemini 2.5 Flash supports a 1,048,576-token context window (roughly 1M tokens) and can generate up to 65,535 output tokens per response. Input and output share this budget, so long inputs reduce room for generation. The large window supports long documents, multi-file inputs, and extended conversation history in a single request.

Yes. Gemini 2.5 Flash was Google's first Flash model with thinking capabilities. Reasoning is configurable through a thinking budget, so you can turn it off for fast, cheap responses or enable it for harder tasks. This lets one endpoint serve both latency-sensitive and reasoning-heavy workloads while controlling cost and latency per request.

Gemini 2.5 Flash accepts text, code, images, audio, video, and documents such as PDFs as input, and returns text only. It does not generate images, audio, or video. This makes it well suited to multimodal understanding and extraction tasks rather than media generation.

Independent testing by Artificial Analysis measures output around 223 tokens per second with time-to-first-token near 0.49 seconds on Google AI Studio in non-reasoning mode, ranking it among the fastest production models. Note that Flash is verbose and often generates more tokens per task, which can offset the per-token speed advantage in end-to-end latency.

Choose Gemini 2.5 Flash when throughput, low latency, and cost efficiency matter most, such as high-volume chat, extraction, and agentic loops. Choose Gemini 2.5 Pro when a task needs maximum reasoning quality or the strongest long-context and coding accuracy. Many teams default to Flash and escalate only the hardest requests to Pro.

* Benchmark data source: Artificial Analysis artificialanalysis.ai