Google's hybrid-reasoning Flash preview with stronger agentic tool use and lower thinking-token cost for high-volume workloads.
Output Speed *
Intelligence Index *
Context Window *
Input price
Output price
Performance
Where Gemini 2.5 Flash Preview 09-2025 Gains Ground
Benchmarks
Independent Benchmarks: Intelligence and Speed at the Flash Tier
Output Speed
Intelligence Index
Technical Specifications
What the model supports
Limitations & Trade-offs
Best-Fit Workloads
Where this model earns its place
Agentic tool-use pipelines
The headline improvement in this release is agentic tool use, with Google reporting better performance on complex, multi-step applications and a five-point SWE-Bench Verified gain. Combined with function calling and structured outputs, it works as an agent backbone for tool-calling loops, orchestration, and autonomous task execution where Flash-tier speed keeps per-step latency low.
Coding assistants and bug-fix automation
A 54% SWE-Bench Verified score reflects real-world software engineering tasks like bug fixes and feature implementation across production codebases. That makes the model a practical fit for code-repair services, PR assistants, and IDE integrations that need solid coding quality without paying Pro-tier rates, provided you chunk output around the 65K cap.
Long-context multimodal analysis
With a ~1M-token context window and text, image, audio and video input, the model suits document understanding, transcript analysis, and mixed-media RAG over large inputs. The wide context lets you pass entire documents or media transcripts in one request; just remember output is text-only and capped at 65K tokens per response.
High-throughput reasoning at scale
The ~24% reduction in output tokens with thinking enabled lowers cost and latency per reasoning-intensive request, and Artificial Analysis places the model among the faster options for its intelligence level. This makes it well-suited to high-volume classification, extraction, and analysis pipelines where reasoning quality matters but per-request cost must stay controlled.