Google
•
Gemini 2.5 Flash-Lite Preview (Sep '25)
•
Released 
September 2025

Google
Gemini 2.5 Flash-Lite Preview (Sep '25)

Google's fastest, lowest-cost Gemini 2.5 model for high-volume classification, extraction and routing at ultra-low latency.

Modality:
Text
Image
Audio
Video
model ID
google/gemini-2.5-flash-lite-preview-09-2025

Output Speed *

N/A
tok/s

Intelligence Index *

13.1
/ 100

Context Window *

1048576
tokens

Input price

0.6
Anytoken

Output price

2.4
Anytoken
Gemini 2.5 Flash Lite Preview 09-2025: High-Throughput Inference at the Lowest Cost Tier Gemini 2.5 Flash Lite Preview 09-2025 is Google's most cost-efficient and lowest-latency model in the Gemini 2.5 family. This preview refines the Flash-Lite line with reduced verbosity, better instruction following and stronger multimodal handling. It accepts text, image, audio, video and PDF input, returns text, and carries a 1M-token context window. Reasoning is off by default to maximize speed, but a controllable thinking budget can be enabled per request. It suits high-volume classification, extraction and routing where throughput and unit economics matter more than frontier intelligence. Integrate Gemini 2.5 Flash Lite Preview 09-2025 via the AnyAPI.ai API

Performance

Where Flash-Lite 09-2025 Earns Its Place: Speed and Token Efficiency

This preview's strongest characteristic is raw generation speed paired with reduced output volume. Artificial Analysis measured Flash-Lite Preview 09-2025 (reasoning) at roughly 887 output tokens per second on AI Studio, the fastest proprietary model in their harness. Google reports a roughly 50% reduction in output tokens versus the earlier version, lowering deployment cost in high-throughput applications. For production, that combination means lower cost per completed task and tighter latency budgets, making the model a strong default for large-scale pipelines where volume, not frontier reasoning, drives spend.

Benchmarks

Independent Intelligence and Speed Data from Artificial Analysis

On the Artificial Analysis Intelligence Index, <cite index="20-3,20-4">Gemini 2.5 Flash-Lite Preview 09-2025 scores 48 in reasoning mode, an 8-point uplift over the June release, and 42 in non-reasoning mode, a 12-point gain over the July version.</cite> On throughput, <cite index="21-1,21-6">Artificial Analysis reports an output speed of 734.8 tokens per second, with a higher-than-average time to first token of 5.05 seconds in reasoning mode.</cite> The intelligence jump is meaningful for a lightweight model, but scores remain below frontier reasoning models—confirming its role as a fast, economical workhorse rather than a top-tier reasoner.

Output Speed

*
N/A
tok/s

Intelligence Index

*
13.1
/ 100

Technical Specifications

What the model supports

The model pairs a 1,048,576-token context window with a maximum text output of 65,536 tokens. <cite index="30-1,30-2">Inputs span text, image, video, audio and PDF, and it is positioned for high-volume classification, simple data extraction and extremely low-latency applications where budget and speed are the primary constraints.</cite> The most production-relevant detail is reasoning: <cite index="23-6,23-7">thinking is disabled by default to prioritize speed, but developers can enable it via the reasoning parameter to trade cost for intelligence.</cite> This means the same endpoint can serve cheap, fast defaults and occasional deeper reasoning without switching models.
Verified Specifications — 
Gemini 2.5 Flash-Lite Preview (Sep '25)
*
Input modalities
Text
Image
Audio
Video
Context window
1048576
 tokens
Reasoning
Yes
Knowledge cutoff
September 2025
Pricing (standard)
0.6
 AnyTokens in
 / 
2.4
 AnyTokens out

Limitations & Trade-offs

Where Gemini 2.5 Flash-Lite Preview (Sep '25) falls short

1
Preview and deprecation status. This is explicitly a preview build, and Google stated these releases were not intended to graduate to a stable version. The endpoint has since been shut down as of March 31, 2026, with migration recommended to Gemini 3.1 Flash-Lite. Teams pinning this exact model string should plan migration paths rather than treating it as a long-term stable dependency.
2
Below-frontier intelligence. With an Artificial Analysis Intelligence Index of 48 in reasoning mode, Flash-Lite trails frontier reasoning models and its own sibling Flash. For complex coding, deep multi-step reasoning or nuanced analysis, a higher-tier model is preferable. Flash-Lite is engineered for volume and speed, not for tasks where accuracy on hard problems is the priority.
3
Reasoning adds latency and cost. Thinking is disabled by default; enabling it raises token consumption and latency because reasoning tokens are billed. On top of this, independent measurement showed a comparatively high time to first token (around 5 seconds) in reasoning mode. For interactive chat or voice where responsiveness matters, keep thinking off or budget it tightly and validate perceived latency on your own workload.
4
Text-only output. The model ingests images, audio, video and PDFs but returns text only. Workloads that need generated images, audio output or non-text artifacts require a different model. If your pipeline expects structured non-text results, plan for text-plus-schema outputs or route those steps elsewhere.

Best-Fit Workloads

Where this model earns its place

01

High-volume classification and routing

‍
Google positions Flash-Lite for high-volume classification, data extraction and low-latency routing. Its combination of very high output speed and reduced verbosity makes per-request cost predictable at scale. This fits content moderation triage, intent classification, ticket routing and pre-filtering stages ahead of a more capable model. Keep thinking disabled to hold latency and cost at their lowest.

02

Large-batch document and multimodal extraction

‍
With a 1M-token context window and input support for text, images, audio, video and PDFs, the model handles bulk extraction across long documents and mixed media. Suitable for parsing invoices, transcribing and summarizing audio, or pulling structured fields from PDFs at volume. Output is text only, so pair it with JSON schema structured outputs when downstream systems need machine-readable results.

03

Latency-sensitive, high-throughput generation

‍
Independent benchmarks rank this preview among the fastest proprietary models measured, at roughly 887 tokens per second in one setup. That throughput suits streaming responses, real-time suggestions and other high-frequency generation where every millisecond and every token counts. Validate time to first token for your prompt sizes, since latency rises when reasoning is enabled.

04

Cost-sensitive agent sub-tasks

‍
Within a larger agent stack, Flash-Lite is well suited to the cheap, frequent steps—tool selection, short summarization, formatting and light transformation—while a stronger model handles hard reasoning. Its optional thinking budget lets you selectively raise intelligence for specific sub-tasks without switching models, keeping most of the pipeline on the lowest-cost tier.

Pricing in anytokens via AnyAPI
Input
0.6
₳
Output
2.4
₳
Cache write
—
₳
Cache read
—
₳

Integration

Access Gemini 2.5 Flash-Lite Preview (Sep '25) via AnyAPI.ai

Access Gemini 2.5 Flash-Lite Preview (Sep '25) through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate Gemini 2.5 Flash-Lite Preview (Sep '25) without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access Gemini 2.5 Flash-Lite Preview (Sep '25) and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test Gemini 2.5 Flash-Lite Preview (Sep '25) against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use Gemini 2.5 Flash-Lite Preview (Sep '25) from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use Gemini 2.5 Flash-Lite Preview (Sep '25) for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

The model has a 1,048,576-token (about 1M) context window and produces up to 65,536 output tokens. Context is shared across all input, including multimodal content such as images, audio, video and PDFs. Output is text only.

Yes. It is a thinking model, but reasoning is disabled by default to prioritize speed and cost. Developers can enable it through the reasoning parameter and set a controllable thinking budget, paying for reasoning tokens only when used. Enabling thinking increases latency and token consumption.

Gemini 2.5 Flash Lite Preview 09-2025 accepts text, images, audio, video and PDF/document input and returns text output. It also supports function calling, structured JSON outputs, and native tools including Grounding with Google Search, code execution and URL context.

This was a preview release and, per Google's documentation, it has been shut down as of March 31, 2026, with migration recommended to Gemini 3.1 Flash-Lite. Google indicated these previews were not intended to graduate to a stable version. Teams should plan a migration path when relying on this exact model string.

Independent benchmarking by Artificial Analysis measured it among the fastest proprietary models tested, at roughly 887 output tokens per second on AI Studio in one setup, with about 734.8 tokens per second in their reasoning-mode figures. Google also reported it is roughly 40% faster than its July 2025 release.

* Benchmark data source: Artificial Analysis artificialanalysis.ai