Google
Gemini 2.5 Flash Image
Released 
August 2025

Google
Gemini 2.5 Flash Image

Native image generation and conversational editing with strong character consistency and fast, high-volume creative iteration.

Modality:
Text
Image
PDF
model ID
google/gemini-2.5-flash-image

Output Speed

N/A
tok/s

Intelligence Index

14.2
/ 100

Context Window

32768
tokens

Input price

1.8
Anytoken

Output price

15
Anytoken
Nano Banana: Conversational Image Generation and Editing With Character Consistency Gemini 2.5 Flash Image, nicknamed Nano Banana, is Google's native image generation and editing model built on the speed and cost profile of Gemini 2.5 Flash. It accepts text and image input and returns text and image output, letting you generate images, edit existing ones, blend multiple references, and maintain the same character or object across turns using natural-language prompts. Positioned as a fast, high-volume creative model rather than a reasoning engine, it fits marketing asset production, product imagery, and low-latency editing workflows where iteration speed and consistency matter more than deep reasoning. Start generating and editing images with Nano Banana via the AnyAPI.ai API.

Performance

Where Nano Banana Leads: Editing Consistency and Iteration Speed

Nano Banana's strongest area is image editing and character consistency: it keeps the same subject recognizable across multiple edits and blends multiple reference images with natural-language instructions. In blind pre-release testing on LMArena under its codename, it ranked #1 on the Image Edit leaderboard with the largest Elo lead recorded there at the time. For production, this means predictable identity preservation across generations—valuable for brand assets, product shots, and multi-turn creative flows—without the fine-tuning overhead earlier pipelines required. It targets fast, cost-efficient, high-volume creative work rather than reasoning-heavy tasks.

Benchmarks

How Nano Banana Ranked in Human-Preference Testing

Independent human-preference evaluation on LMArena placed Nano Banana at the top of image editing during pre-release testing under its codename, driving millions of community votes and the largest Elo lead in that arena's history at the time. It also topped text-to-image preference rankings in late August 2025. These are human-preference measurements of realism and instruction adherence, not intelligence or coding scores—no reasoning benchmarks are published for this model, since it is an image model rather than a text-reasoning model. Interpret the results as evidence of editing quality and prompt fidelity, not general capability.

Output Speed

N/A
tok/s

Intelligence Index

14.2
/ 100

MMLU

Broad world knowledge and problem-solving
81
%

GPQA

PhD-level scientific reasoning across physics, biology, chemistry.
68
%

HLE

Adherence to multi-step structured instructions.
5
%

LiveCodeBench

Tool-calling reliability in long agentic loops.
50
%

Technical Specifications

What the model supports

Gemini 2.5 Flash Image accepts text and image input and returns both text and image output. Input and output token limits are each 32,768. In practice each generated 1024x1024 image consumes roughly 1290 output tokens, so the token ceiling is generous for image workloads. The model supports up to 3 input images per prompt, up to 10 output images, and 10 aspect ratios from 21:9 to 9:16. All generated or edited images embed an invisible SynthID watermark. This is a creative-output model, not a text-reasoning model—plan a separate model for reasoning-heavy tasks.
Verified Specifications — 
Gemini 2.5 Flash Image
Input modalities
Text
Image
PDF
output modalities
Text
Image
Context window
32768
 tokens
Maximum output tokens
32768
Reasoning
Yes
Knowledge cutoff
August 2025
Pricing (standard)
1.8
 AnyTokens in
 / 
15
 AnyTokens out

Limitations & Trade-offs

Where Gemini 2.5 Flash Image falls short

1
No reasoning or text-analysis capability. Nano Banana is an image generation and editing model with no published reasoning benchmarks and no extended-thinking mode. It cannot substitute for a text model on coding, agentic tasks, or long-document analysis. Teams needing both visuals and reasoning must pair it with a separate text-focused model rather than routing everything through Nano Banana.
2
Limited input images per prompt. On Vertex AI the model accepts a maximum of 3 input images per prompt. Workflows that need to fuse many reference images into a single composition are constrained here; newer Nano Banana models support far more reference images. If your creative pipeline depends on blending large numbers of source images, this generation may be too restrictive.
3
Mandatory SynthID watermarking. Every image generated or edited by the model embeds an invisible SynthID watermark for provenance. This is unavoidable and by design. It's beneficial for transparency and AI-content identification, but teams that require unmarked outputs for specific downstream uses should account for it before standardizing on this model.
4
Superseded by newer Nano Banana models. Google now classifies Gemini 2.5 Flash Image as a legacy model and recommends newer Nano Banana releases for better quality, faster generation, and lower cost. It remains available and reliable, but new projects planning long production lifetimes should evaluate whether a current-generation image model is the safer default.

Best-Fit Workloads

Where this model earns its place

01

Conversational image editing


The model's core strength is multi-turn editing driven by natural language—changing backgrounds, adjusting lighting, adding or removing objects—while keeping the rest of the image stable. Its top LMArena image-editing ranking reflects this. It fits editing tools and creative apps where users refine an image across several steps rather than regenerating from scratch each time.

02

Brand and character-consistent asset generation


Nano Banana preserves the same character, product, or style across multiple generations without fine-tuning. This makes it well suited to producing consistent brand assets, showing a product from multiple angles, or placing a recurring character in different scenes—use cases where identity drift between images would break the workflow.

03

High-volume marketing and social imagery


Built on the speed and cost profile of Gemini 2.5 Flash and supporting 10 aspect ratios from cinematic 21:9 to vertical 9:16, the model targets fast, high-throughput creative production. It fits marketing pipelines generating many variations across social, ad, and product formats where iteration speed and per-image cost matter more than deep reasoning.

04

Multi-image fusion


The model blends multiple input images into a single cohesive composition through natural-language instructions—for example combining a product cutout with a new background. This suits composition and mockup workflows, though the 3-input-image limit on Vertex AI caps how many sources can be fused in one prompt.

Pricing in anytokens via AnyAPI
Input
1.8
Output
15
Cache write
Cache read

Integration

Access Gemini 2.5 Flash Image via AnyAPI.ai

Access Gemini 2.5 Flash Image through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate Gemini 2.5 Flash Image without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access Gemini 2.5 Flash Image and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test Gemini 2.5 Flash Image against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use Gemini 2.5 Flash Image from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use Gemini 2.5 Flash Image for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

Gemini 2.5 Flash Image, nicknamed Nano Banana, is Google's native image generation and editing model released August 26, 2025. It accepts text and image input and produces text and image output, supporting generation, conversational editing, multi-image fusion, and character consistency across turns. It's built on the speed and cost profile of Gemini 2.5 Flash.

No. Gemini 2.5 Flash Image is an image generation and editing model with no reasoning or extended-thinking mode and no published reasoning benchmarks. It is not designed for coding, agentic tasks, or long-document analysis. Pair it with a separate text-focused model when you need reasoning alongside visuals.

The model has a 32,768-token input context window and a 32,768-token maximum output. Each generated 1024x1024 image consumes roughly 1290 output tokens. On Vertex AI it accepts up to 3 input images per prompt and can return up to 10 output images.

Yes. Every image generated or edited with Gemini 2.5 Flash Image embeds an invisible SynthID watermark so the output can be identified as AI-generated or AI-edited. This is applied by default and cannot be disabled, supporting content provenance and transparency.

In blind human-preference testing on LMArena under its codename, Nano Banana ranked #1 on the Image Edit leaderboard with the largest Elo lead recorded there at the time, and topped text-to-image preference rankings in late August 2025. These measure editing quality and instruction adherence, not general intelligence.