Google
•
Gemini 2.5 Flash-Lite Preview (Sep '25)
•
Released 
September 2025

Google
Gemini 2.5 Flash-Lite Preview (Sep '25)

Google's hybrid-reasoning Flash preview with stronger agentic tool use and lower thinking-token cost for high-volume workloads.

Modality:
Text
Image
Video
model ID
google/gemini-2.5-flash-lite-preview-09-2025

Output Speed *

N/A
tok/s

Intelligence Index *

13.1
/ 100

Context Window *

1048576
tokens

Input price

1.8
Anytoken

Output price

15
Anytoken
Gemini 2.5 Flash Preview 09-2025: Faster Agentic Tool Use With Cheaper Thinking Gemini 2.5 Flash Preview 09-2025 is Google's September 2025 checkpoint of the 2.5 Flash tier, a hybrid model that can answer directly or engage extended thinking. It sits between the lightweight Flash-Lite and higher-cost Pro models as a balanced, high-throughput workhorse. This preview targets two areas: improved agentic tool use across multi-step tasks, and greater token efficiency when thinking is enabled, so reasoning-heavy prompts finish with fewer output tokens. It suits agent loops, coding assistants, and multimodal pipelines that need strong quality at Flash-tier speed and cost. Start building with Gemini 2.5 Flash Preview 09-2025 via the AnyAPI.ai API.

Performance

Where Gemini 2.5 Flash Preview 09-2025 Gains Ground

This preview's strongest gains are in agentic tool use and reasoning efficiency. Google reports SWE-Bench Verified rising from 48.9% to 54%, a real-world software engineering benchmark covering bug fixes and feature work. Artificial Analysis independently measured 54 in reasoning mode on its Intelligence Index, a three-point lift over the May 2025 Flash. Crucially, the model produces fewer output tokens with thinking enabled, roughly a 24% reduction per Google. For production agents and coding loops, that means comparable quality with lower latency and cost per successful task, which matters most in high-volume, multi-step pipelines.

Benchmarks

Independent Benchmarks: Intelligence and Speed at the Flash Tier

Artificial Analysis benchmarked this preview at 54 on its Intelligence Index in reasoning mode and 47 in non-reasoning mode, representing three- and eight-point gains over the May 2025 Gemini 2.5 Flash. Artificial Analysis noted Google's Flash models offer some of the fastest output speeds relative to models of equivalent intelligence. Vals AI's evaluation of the thinking variant found meaningful improvements over the prior version, including a large jump on GPQA Diamond and gains on Terminal-Bench. Treat these as third-party measurements taken under specific reasoning settings; validate on your own workload before committing in production.

Output Speed

*
N/A
tok/s

Intelligence Index

*
13.1
/ 100

Technical Specifications

What the model supports

The model accepts text, image, audio and video input and returns text only, with a ~1M-token context window and up to 65,535 output tokens per request. It is a hybrid reasoning model: thinking can be enabled and budgeted via a reasoning-token parameter. The two most production-relevant limits are the fixed 65K output cap, which constrains very long single-response generation, and its preview status, which carries deprecation notices and rate-limit variability. Function calling and structured JSON outputs are supported, making it usable as an agent backbone rather than a chat-only endpoint.
Verified Specifications — 
Gemini 2.5 Flash-Lite Preview (Sep '25)
*
Input modalities
Text
Image
Video
Context window
1048576
 tokens
Maximum output tokens
65536
Reasoning
Yes
Knowledge cutoff
September 2025
Pricing (standard)
1.8
 AnyTokens in
 / 
15
 AnyTokens out

Limitations & Trade-offs

Where Gemini 2.5 Flash-Lite Preview (Sep '25) falls short

1
Preview lifecycle risk. This is a preview checkpoint, not a stable release. Google's documentation notes preview models come with more restrictive rate limits and at least two weeks' deprecation notice, and this specific endpoint was later shut down. For production systems that cannot re-validate against a moving target, the stable Gemini 2.5 Flash or a current generally available Flash model is a safer default.
2
65K output ceiling. Maximum output is capped at 65,535 tokens per request, well below the ~1M-token input window. Workloads that need very long single responses — full-length document generation or exhaustive code files in one call — will hit this ceiling and require chunking or multi-turn strategies. This is a hard limit, not a tuning parameter.
3
Thinking-token billing. Reasoning is billed, and output tokens on the Flash tier cost materially more than input tokens. Even with the release's ~24% token-efficiency improvement, reasoning-heavy prompts at high volume can accumulate cost. Budget the thinking parameter deliberately, or disable thinking for straightforward tasks where the non-reasoning path is sufficient.
4
Text-only output. Despite accepting text, image, audio and video input, the model returns text only. Applications that need generated images, audio, or speech must route those to dedicated Gemini image or audio models. Do not assume this endpoint inherits the family's image-generation capabilities.

Best-Fit Workloads

Where this model earns its place

01

Agentic tool-use pipelines

‍
The headline improvement in this release is agentic tool use, with Google reporting better performance on complex, multi-step applications and a five-point SWE-Bench Verified gain. Combined with function calling and structured outputs, it works as an agent backbone for tool-calling loops, orchestration, and autonomous task execution where Flash-tier speed keeps per-step latency low.

02

Coding assistants and bug-fix automation

‍
A 54% SWE-Bench Verified score reflects real-world software engineering tasks like bug fixes and feature implementation across production codebases. That makes the model a practical fit for code-repair services, PR assistants, and IDE integrations that need solid coding quality without paying Pro-tier rates, provided you chunk output around the 65K cap.

03

Long-context multimodal analysis

‍
With a ~1M-token context window and text, image, audio and video input, the model suits document understanding, transcript analysis, and mixed-media RAG over large inputs. The wide context lets you pass entire documents or media transcripts in one request; just remember output is text-only and capped at 65K tokens per response.

04

High-throughput reasoning at scale

‍
The ~24% reduction in output tokens with thinking enabled lowers cost and latency per reasoning-intensive request, and Artificial Analysis places the model among the faster options for its intelligence level. This makes it well-suited to high-volume classification, extraction, and analysis pipelines where reasoning quality matters but per-request cost must stay controlled.

Pricing in anytokens via AnyAPI
Input
1.8
₳
Output
15
₳
Cache write
—
₳
Cache read
—
₳

Integration

Access Gemini 2.5 Flash-Lite Preview (Sep '25) via AnyAPI.ai

Access Gemini 2.5 Flash-Lite Preview (Sep '25) through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate Gemini 2.5 Flash-Lite Preview (Sep '25) without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access Gemini 2.5 Flash-Lite Preview (Sep '25) and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test Gemini 2.5 Flash-Lite Preview (Sep '25) against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use Gemini 2.5 Flash-Lite Preview (Sep '25) from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use Gemini 2.5 Flash-Lite Preview (Sep '25) for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

The model supports a context window of 1,048,576 tokens (~1M), shared across text, image, audio and video input. Maximum output is separate and capped at 65,535 tokens per request. The large input window suits long documents and multimodal RAG, but long single responses must be chunked around the output ceiling.

Yes. It is a hybrid model with built-in thinking capabilities that can be enabled and budgeted through a reasoning-token parameter. This release notably reduces output tokens with thinking on, cutting cost and latency on reasoning-intensive prompts while maintaining quality, per Google's own reporting.

The September 2025 preview improves agentic tool use, raises SWE-Bench Verified from 48.9% to 54%, and reduces thinking-mode output tokens by roughly 24% versus the May 2025 stable Flash. The trade-off is preview lifecycle risk: preview endpoints carry deprecation notices and more restrictive rate limits.

The two biggest constraints are its preview status, which carries deprecation and rate-limit risk, and the 65,535-token output cap. Output is also text-only despite multimodal input. Teams needing long-term stability should consider a generally available Flash model instead.

It accepts text, image, audio and video as input and returns text only. Function calling and structured JSON outputs are supported, making it suitable as an agent backbone. For generated images or audio, route to dedicated Gemini image or audio models rather than this endpoint.

* Benchmark data source: Artificial Analysis artificialanalysis.ai