Google's fastest, lowest-cost Gemini 2.5 model for high-volume classification, extraction and routing at ultra-low latency.
Output Speed *
Intelligence Index *
Context Window *
Input price
Output price
Performance
Where Flash-Lite 09-2025 Earns Its Place: Speed and Token Efficiency
Benchmarks
Independent Intelligence and Speed Data from Artificial Analysis
Output Speed
Intelligence Index
Technical Specifications
What the model supports
Limitations & Trade-offs
Best-Fit Workloads
Where this model earns its place
High-volume classification and routing
Google positions Flash-Lite for high-volume classification, data extraction and low-latency routing. Its combination of very high output speed and reduced verbosity makes per-request cost predictable at scale. This fits content moderation triage, intent classification, ticket routing and pre-filtering stages ahead of a more capable model. Keep thinking disabled to hold latency and cost at their lowest.
Large-batch document and multimodal extraction
With a 1M-token context window and input support for text, images, audio, video and PDFs, the model handles bulk extraction across long documents and mixed media. Suitable for parsing invoices, transcribing and summarizing audio, or pulling structured fields from PDFs at volume. Output is text only, so pair it with JSON schema structured outputs when downstream systems need machine-readable results.
Latency-sensitive, high-throughput generation
Independent benchmarks rank this preview among the fastest proprietary models measured, at roughly 887 tokens per second in one setup. That throughput suits streaming responses, real-time suggestions and other high-frequency generation where every millisecond and every token counts. Validate time to first token for your prompt sizes, since latency rises when reasoning is enabled.
Cost-sensitive agent sub-tasks
Within a larger agent stack, Flash-Lite is well suited to the cheap, frequent steps—tool selection, short summarization, formatting and light transformation—while a stronger model handles hard reasoning. Its optional thinking budget lets you selectively raise intelligence for specific sub-tasks without switching models, keeping most of the pipeline on the lowest-cost tier.