Google
•
Gemini 2.0 Flash (Feb '25)
•
Released 
February 2025

Google
Gemini 2.0 Flash (Feb '25)

Google's high-volume workhorse model combining fast multimodal inference, a 1M-token context window, and native tool use.

Modality:
Text
Image
Audio
Video
PDF
model ID
google/gemini-2.0-flash-001

Output Speed *

0.00
tok/s

Intelligence Index *

8.9
/ 100

Context Window *

1000000
tokens

Input price

0.9
Anytoken

Output price

3.6
Anytoken
Gemini 2.0 Flash: Fast Multimodal Inference at Million-Token Scale Gemini 2.0 Flash is Google DeepMind's workhorse model for the Gemini 2.0 family, positioned between the cost-optimized Flash-Lite and the more capable Pro tier. It processes text, images, audio, video, and PDFs against a 1M-token context window, returns text, and includes native tool use and function calling. Its defining trait is throughput: low time to first token and high output speed at a competitive price. This makes it well suited to high-frequency, high-volume production tasks such as document processing, multimodal extraction, and agentic pipelines where latency and cost matter more than frontier reasoning. Access Gemini-family models and supported successors through one API on AnyAPI.ai.

Performance

Where Gemini 2.0 Flash Earns Its Place: Throughput and Multimodal Coverage

Gemini 2.0 Flash is built for volume rather than frontier reasoning. Independent benchmarking measured roughly 169.5 output tokens per second and around 0.53s to first token, placing it among faster non-reasoning models at its price point. It scores 12 on the Artificial Analysis Intelligence Index, above the average for comparable models. In practice, that combination means you can run high-frequency multimodal tasks—image, PDF, audio and video understanding—with predictable latency. The production consequence is straightforward: for interactive apps and large batch pipelines, response speed and cost efficiency usually matter more than a few points of benchmark intelligence, and this model optimizes for exactly that.

Benchmarks

Gemini 2.0 Flash Benchmarks: Solid Non-Reasoning Scores, Fast Delivery

On independent evaluations, Gemini 2.0 Flash posts an Artificial Analysis Intelligence Index of 12, above the comparable-model average of 11 for non-reasoning models. Third-party testing reports MMLU-Pro around 77.9 and GPQA Diamond around 62.3, with LiveCodeBench near 33.4—competent mid-tier results rather than leadership numbers. The stronger story is delivery: measured output speed of roughly 169.5 tokens per second and time to first token near 0.53s. Interpretation matters here: these are independent measurements, not official Google specifications, and reasoning-tuned models will outperform it on hard multi-step tasks while running slower and costing more.

Output Speed

*
0.00
tok/s

Intelligence Index

*
8.9
/ 100

MMLU *

Broad world knowledge and problem-solving
78
%

GPQA *

PhD-level scientific reasoning across physics, biology, chemistry.
62
%

HLE *

Adherence to multi-step structured instructions.
4
%

LiveCodeBench *

Tool-calling reliability in long agentic loops.
33
%

Technical Specifications

What the model supports

Gemini 2.0 Flash accepts text, images, audio, video, and PDF documents as input and returns text; it is not a text-to-image or audio-output model in general API use. It exposes a 1,048,576-token context window but caps output at 8,192 tokens by default—the two limits are independent, so long inputs do not extend generation length. Native function calling and JSON-schema structured outputs make it viable for agentic and data-extraction workloads. The large context enables whole-document and long-video analysis in a single request, but the modest output ceiling means long-form generation must be chunked.
Verified Specifications — 
Gemini 2.0 Flash (Feb '25)
*
Input modalities
Text
Image
Audio
Video
PDF
output modalities
Text
Context window
1000000
 tokens
Maximum output tokens
8192
Reasoning
Yes
Knowledge cutoff
February 2025
Pricing (standard)
0.9
 AnyTokens in
 / 
3.6
 AnyTokens out

Quickstart

Sample code for Gemini 2.0 Flash (Feb '25)

import requests

url = "https://api.anyapi.ai/v1/chat/completions"

payload = {
    "stream": False,
    "tool_choice": "auto",
    "logprobs": False,
    "model": "gemini-2.0-flash-001",
    "messages": [
        {
            "content": [
                {
                    "type": "text",
                    "text": "Hello"
                },
                {
                    "image_url": {
                        "detail": "auto",
                        "url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"
                    },
                    "type": "image_url"
                }
            ],
            "role": "user"
        }
    ]
}
headers = {
    "Authorization": "Bearer AnyAPI_API_KEY",
    "Content-Type": "application/json"
}

response = requests.post(url, json=payload, headers=headers)

print(response.json())
import requests url = "https://api.anyapi.ai/v1/chat/completions" payload = { "stream": False, "tool_choice": "auto", "logprobs": False, "model": "gemini-2.0-flash-001", "messages": [ { "content": [ { "type": "text", "text": "Hello" }, { "image_url": { "detail": "auto", "url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg" }, "type": "image_url" } ], "role": "user" } ] } headers = { "Authorization": "Bearer AnyAPI_API_KEY", "Content-Type": "application/json" } response = requests.post(url, json=payload, headers=headers) print(response.json())
View docs
Copy
Code is copied
const url = 'https://api.anyapi.ai/v1/chat/completions';
const options = {
  method: 'POST',
  headers: {Authorization: 'Bearer AnyAPI_API_KEY', 'Content-Type': 'application/json'},
  body: '{"stream":false,"tool_choice":"auto","logprobs":false,"model":"gemini-2.0-flash-001","messages":[{"content":[{"type":"text","text":"Hello"},{"image_url":{"detail":"auto","url":"https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"},"type":"image_url"}],"role":"user"}]}'
};

try {
  const response = await fetch(url, options);
  const data = await response.json();
  console.log(data);
} catch (error) {
  console.error(error);
}
const url = 'https://api.anyapi.ai/v1/chat/completions'; const options = { method: 'POST', headers: {Authorization: 'Bearer AnyAPI_API_KEY', 'Content-Type': 'application/json'}, body: '{"stream":false,"tool_choice":"auto","logprobs":false,"model":"gemini-2.0-flash-001","messages":[{"content":[{"type":"text","text":"Hello"},{"image_url":{"detail":"auto","url":"https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"},"type":"image_url"}],"role":"user"}]}' }; try { const response = await fetch(url, options); const data = await response.json(); console.log(data); } catch (error) { console.error(error); }
View docs
Copy
Code is copied
curl --request POST \
  --url https://api.anyapi.ai/v1/chat/completions \
  --header 'Authorization: Bearer AnyAPI_API_KEY' \
  --header 'Content-Type: application/json' \
  --data '{
  "stream": false,
  "tool_choice": "auto",
  "logprobs": false,
  "model": "gemini-2.0-flash-001",
  "messages": [
    {
      "content": [
        {
          "type": "text",
          "text": "Hello"
        },
        {
          "image_url": {
            "detail": "auto",
            "url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"
          },
          "type": "image_url"
        }
      ],
      "role": "user"
    }
  ]
}'
curl --request POST \ --url https://api.anyapi.ai/v1/chat/completions \ --header 'Authorization: Bearer AnyAPI_API_KEY' \ --header 'Content-Type: application/json' \ --data '{ "stream": false, "tool_choice": "auto", "logprobs": false, "model": "gemini-2.0-flash-001", "messages": [ { "content": [ { "type": "text", "text": "Hello" }, { "image_url": { "detail": "auto", "url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg" }, "type": "image_url" } ], "role": "user" } ] }'
View docs
Copy
Code is copied
View docs
Code examples coming soon...

Comparison

Gemini 2.0 Flash vs Gemini 2.5 Flash: What Changes on Upgrade?

These are the natural comparison because Gemini 2.5 Flash is Google's designated migration target after Gemini 2.0 Flash's June 2026 shutdown. Both are Flash-tier workhorse models with a 1M-token context window, multimodal input, function calling, and structured outputs. The key difference is capability and design: Gemini 2.5 Flash adds controllable thinking (an extended-reasoning budget) and scores higher on independent intelligence benchmarks—14 versus 12 on the Artificial Analysis Index. It also generates faster in measured tests. The practical decision is whether you need optional reasoning and higher intelligence, or a simpler, lower-cost non-reasoning model.

Dimension
Gemini 2.0 Flash (Feb '25)
Gemini 2.5 Flash (Non-reasoning)
Context window *
1000000
tokens
1048576
tokens
Output speed *
0.00
tok/s
0.00
tok/s
Intelligence Index *
8.9
9.9
Input pricing
0.9
AnyToken
1.8
AnyToken
Output pricing
3.6
AnyToken
15
AnyToken
Knowledge cutoff *
February 2025
June 2025

Choose Gemini 2.0 Flash when you want a straightforward non-reasoning model for high-volume multimodal and extraction tasks and are optimizing for lower cost per token—though note it is deprecated and should only be considered where already integrated. Choose Gemini 2.5 Flash for any new build: it offers higher benchmark intelligence, optional thinking budgets for harder tasks, faster measured output, and an active support lifecycle. For new production systems, 2.5 Flash is the safer long-term choice; 2.0 Flash mainly matters for understanding legacy deployments.

Limitations & Trade-offs

Where Gemini 2.0 Flash (Feb '25) falls short

1
Deprecated and shut down. Gemini 2.0 Flash reached its shutdown date on June 1, 2026, and Google directs developers to Gemini 2.5 Flash instead. This is the single most important limitation for any production decision: new integrations cannot rely on it, and existing ones must migrate. It matters most for teams evaluating the model today, who should treat this page as reference and select an actively supported Flash-tier successor.
2
No native reasoning mode. Gemini 2.0 Flash is a non-reasoning model that returns direct responses without an extended chain-of-thought budget. On hard multi-step math, agentic planning, or complex coding, it trails reasoning-tuned models, and independent benchmarks place its intelligence in the mid tier. For workloads that depend on deep reasoning, a model with a controllable thinking budget—such as Gemini 2.5 Flash or a dedicated reasoning model—is a better fit.
3
Modest output ceiling. Despite the 1M-token input context, the default maximum output is 8,192 tokens. Long-context input does not translate into long-form generation, so tasks like full-report drafting or large document synthesis must be chunked and stitched. Applications that need very long single-pass outputs will hit this limit and should plan for segmentation or choose a model with a higher output cap.
4
Text-only output in general API use. Gemini 2.0 Flash accepts rich multimodal input but returns text; native image and audio output were staged separately rather than standard in the general text endpoint. If your workload requires generated images or spoken audio as output, this model does not cover it and you will need a dedicated image or text-to-speech model.

Best-Fit Workloads

Where this model earns its place

01

Multimodal document and media understanding

‍
The 1M-token context plus image, PDF, audio, and video input make Gemini 2.0 Flash effective for parsing long documents, transcribing and summarizing recordings, or analyzing lengthy video in a single request. Vertex AI supports large per-prompt file volumes and long video inputs, so entire files can be processed without heavy chunking. Best for extraction and comprehension pipelines where fast turnaround matters; remember output is capped, so summaries scale better than full regeneration.

02

High-volume, low-latency inference
‍

Measured output speed near 169.5 tokens per second and time to first token around 0.53s make this a strong fit for interactive chat, autocomplete, classification, and other high-frequency tasks. Google positions the Flash series as a workhorse for high-volume, high-frequency use at scale. Its concise default style further reduces cost and latency. Production apps serving many concurrent users benefit most, provided they do not require deep reasoning per request.

03

Structured data extraction

‍
With JSON-mode structured outputs backed by OpenAPI-style schemas (and Pydantic/Zod support), Gemini 2.0 Flash reliably converts unstructured text and multimodal input into type-safe records. This suits data-extraction and database-population pipelines where predictable schema adherence matters more than reasoning depth. The large context lets you extract from long source documents in one pass; validate outputs and add bounded retries for robustness.

04

Tool-using agent steps

‍
Native function calling—including automatic function calling through the Python SDK—lets Gemini 2.0 Flash act as a fast executor in agentic workflows, deciding when to invoke tools and formatting their inputs. Google designed Gemini 2.0 for the agentic era with built-in tool use. It fits the high-frequency, lower-complexity steps of a pipeline; route genuinely hard planning to a reasoning-capable model within the same architecture.

Pricing in anytokens via AnyAPI
Input
0.9
₳
Output
3.6
₳
Cache write
—
₳
Cache read
—
₳

Integration

Access Gemini 2.0 Flash (Feb '25) via AnyAPI.ai

Access Gemini 2.0 Flash (Feb '25) through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate Gemini 2.0 Flash (Feb '25) without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access Gemini 2.0 Flash (Feb '25) and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test Gemini 2.0 Flash (Feb '25) against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use Gemini 2.0 Flash (Feb '25) from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use Gemini 2.0 Flash (Feb '25) for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

Gemini 2.0 Flash supports a context window of 1,048,576 tokens (about 1M), covering text, images, audio, video, and PDF input in a single request. Its maximum output is separate and capped at 8,192 tokens by default, so the large context applies to input, not generation length.

No. Google deprecated Gemini 2.0 Flash and shut it down on June 1, 2026, directing developers to Gemini 2.5 Flash as the migration target. New projects should build on an actively supported Flash-tier model rather than gemini-2.0-flash.

No. Gemini 2.0 Flash is a non-reasoning model that returns direct responses without an extended thinking budget. A separate experimental Flash Thinking variant existed, but the standard model does not include controllable reasoning. For deep multi-step reasoning, Gemini 2.5 Flash or a dedicated reasoning model is a better choice.

Gemini 2.0 Flash accepts text, images, audio, video, and PDF documents as input and returns text output in general API use. It supports native function calling and JSON-schema structured outputs, but does not produce images or audio as standard text-endpoint output.

Independent benchmarking by Artificial Analysis measured Gemini 2.0 Flash at roughly 169.5 output tokens per second with a time to first token near 0.53 seconds—among faster non-reasoning models at its price tier. These are independent measurements, not official Google specifications.

* Benchmark data source: Artificial Analysis artificialanalysis.ai