Google
•
Gemma 3 27B Instruct
•
Released 
March 2025

Google
Gemma 3 27B Instruct

Google's largest open-weight Gemma 3 model, running vision-plus-text workloads on a single GPU with a 128K context.

Modality:
Text
Image
Video
PDF
model ID
google/gemma-3-27b-it

Output Speed *

0.00
tok/s

Intelligence Index *

4.9
/ 100

Context Window *

131072
tokens

Input price

0.48
Anytoken

Output price

0.96
Anytoken
Gemma 3 27B: Multimodal Open Weights That Run on One GPU Gemma 3 27B is the largest model in Google DeepMind's open-weight Gemma 3 family, built from the same research as Gemini. It accepts text and image input and returns text, supports 140+ languages, and exposes a 128K-token context. Its defining trait is capability-per-accelerator: it delivers competitive quality while remaining deployable on a single GPU, including consumer cards with quantization. It suits teams that need self-hosted, license-permissive multimodal inference for document analysis, multilingual generation, and visual question answering without depending on a proprietary API. Start building with the Gemma 3 27B API on AnyAPI.ai.

Performance

Where Gemma 3 27B Earns Its Place: Quality Per Accelerator

Gemma 3 27B's strongest characteristic is delivering solid general and multimodal quality at a size that fits a single accelerator. It scores 67.5 on MMLU-Pro and 89.0 on MATH in Google's technical report, and reached an LMArena Elo around 1338 at release, competing with much larger open models. Because that quality runs on one GPU, teams can self-host multimodal inference without multi-node serving or per-token API dependence. The practical consequence: predictable, license-permissive deployment for document, vision, and multilingual workloads where controlling infrastructure matters more than frontier reasoning scores.

Benchmarks

Gemma 3 27B Benchmarks: Strong Math, Modest Coding and Agents

In Google's Gemma 3 technical report, the instruction-tuned 27B scores 67.5 on MMLU-Pro and 89.0 on MATH, indicating solid knowledge and symbolic reasoning for its size. Coding is more modest: it reaches 29.7 on LiveCodeBench, and independent aggregators place its coding percentile lower among peers. Agentic tool use is a clear weak point, scoring only 6.6 on τ²-bench (Retail) in later Google comparisons. FACTS Grounding of 74.9 contrasts with a low SimpleQA of 10.0, meaning it grounds well on provided documents but is weaker on unaided factual recall.

Output Speed

*
0.00
tok/s

Intelligence Index

*
4.9
/ 100

MMLU *

Broad world knowledge and problem-solving
67
%

GPQA *

PhD-level scientific reasoning across physics, biology, chemistry.
43
%

HLE *

Adherence to multi-step structured instructions.
4
%

LiveCodeBench *

Tool-calling reliability in long agentic loops.
14
%

Technical Specifications

What the model supports

Gemma 3 27B accepts text and image input and produces text output; it is not a reasoning model and has no extended-thinking mode. The 128K-token context and 140+ language coverage are its most production-relevant specifications, enabling long-document and multilingual work in a single request. Note the distinction between context window and output: it processes up to 128K tokens of input but generates far fewer per completion. Open weights under the Gemma license permit commercial use and self-hosting, which is the main reason teams choose it over closed alternatives.
Verified Specifications — 
Gemma 3 27B Instruct
*
Input modalities
Text
Image
Video
PDF
output modalities
Text
Context window
131072
 tokens
Maximum output tokens
16384
Reasoning
No
Knowledge cutoff
March 2025
Pricing (standard)
0.48
 AnyTokens in
 / 
0.96
 AnyTokens out

Quickstart

Sample code for Gemma 3 27B Instruct

import requests

url = "https://api.anyapi.ai/v1/chat/completions"

payload = {
    "stream": False,
    "tool_choice": "auto",
    "logprobs": False,
    "model": "Model_Name",
    "messages": [
        {
            "content": [
                {
                    "type": "text",
                    "text": "Hello"
                },
                {
                    "image_url": {
                        "detail": "auto",
                        "url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"
                    },
                    "type": "image_url"
                }
            ],
            "role": "user"
        }
    ]
}
headers = {
    "Authorization": "Bearer AnyAPI_API_KEY",
    "Content-Type": "application/json"
}

response = requests.post(url, json=payload, headers=headers)

print(response.json())
import requests url = "https://api.anyapi.ai/v1/chat/completions" payload = { "stream": False, "tool_choice": "auto", "logprobs": False, "model": "Model_Name", "messages": [ { "content": [ { "type": "text", "text": "Hello" }, { "image_url": { "detail": "auto", "url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg" }, "type": "image_url" } ], "role": "user" } ] } headers = { "Authorization": "Bearer AnyAPI_API_KEY", "Content-Type": "application/json" } response = requests.post(url, json=payload, headers=headers) print(response.json())
View docs
Copy
Code is copied
const url = 'https://api.anyapi.ai/v1/chat/completions';
const options = {
  method: 'POST',
  headers: {Authorization: 'Bearer AnyAPI_API_KEY', 'Content-Type': 'application/json'},
  body: '{"stream":false,"tool_choice":"auto","logprobs":false,"model":"Model_Name","messages":[{"content":[{"type":"text","text":"Hello"},{"image_url":{"detail":"auto","url":"https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"},"type":"image_url"}],"role":"user"}]}'
};

try {
  const response = await fetch(url, options);
  const data = await response.json();
  console.log(data);
} catch (error) {
  console.error(error);
}
const url = 'https://api.anyapi.ai/v1/chat/completions'; const options = { method: 'POST', headers: {Authorization: 'Bearer AnyAPI_API_KEY', 'Content-Type': 'application/json'}, body: '{"stream":false,"tool_choice":"auto","logprobs":false,"model":"Model_Name","messages":[{"content":[{"type":"text","text":"Hello"},{"image_url":{"detail":"auto","url":"https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"},"type":"image_url"}],"role":"user"}]}' }; try { const response = await fetch(url, options); const data = await response.json(); console.log(data); } catch (error) { console.error(error); }
View docs
Copy
Code is copied
curl --request POST \
  --url https://api.anyapi.ai/v1/chat/completions \
  --header 'Authorization: Bearer AnyAPI_API_KEY' \
  --header 'Content-Type: application/json' \
  --data '{
  "stream": false,
  "tool_choice": "auto",
  "logprobs": false,
  "model": "Model_Name",
  "messages": [
    {
      "content": [
        {
          "type": "text",
          "text": "Hello"
        },
        {
          "image_url": {
            "detail": "auto",
            "url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"
          },
          "type": "image_url"
        }
      ],
      "role": "user"
    }
  ]
}'
curl --request POST \ --url https://api.anyapi.ai/v1/chat/completions \ --header 'Authorization: Bearer AnyAPI_API_KEY' \ --header 'Content-Type: application/json' \ --data '{ "stream": false, "tool_choice": "auto", "logprobs": false, "model": "Model_Name", "messages": [ { "content": [ { "type": "text", "text": "Hello" }, { "image_url": { "detail": "auto", "url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg" }, "type": "image_url" } ], "role": "user" } ] }'
View docs
Copy
Code is copied
View docs
Code examples coming soon...

Limitations & Trade-offs

Where Gemma 3 27B Instruct falls short

1
No reasoning mode. Gemma 3 27B is confirmed as a non-reasoning model that produces direct responses without extended chain-of-thought. On tasks that reward step-by-step deliberation — hard math contests, multi-step planning, complex debugging — it trails models with explicit thinking modes. Its later successor's dramatic gains on AIME and agentic benchmarks illustrate the gap. If your workload depends on deep reasoning or self-correction, choose a dedicated reasoning model instead.
2
Weak agentic tool use. Independent Google comparisons show Gemma 3 27B scoring only 6.6 on τ²-bench (Retail), a measure of multi-step tool calling and error handling. While the model supports function calling and structured outputs on many providers, its reliability in autonomous, multi-turn agent loops is limited. For production agents that call tools, execute steps, and recover from failures, this model is a risky default and a stronger tool-use model is preferable.
3
Modest coding performance. On LiveCodeBench the 27B reaches roughly 29.7, and independent aggregators place its coding percentile low among comparable models. Real-world testing found it useful for small-scale corrections and implementation support but unreliable on large-codebase understanding and complex refactors. It can assist coding, but it is not a strong choice as the primary engine for coding agents or repository-scale generation.
4
Weak unaided factual recall. Gemma 3 27B scores only 10.0 on SimpleQA while scoring 74.9 on FACTS Grounding, meaning it answers well when grounded on supplied documents but is unreliable on open-ended factual questions from memory. Combined with an August 2024 knowledge cutoff and no built-in web search, this makes retrieval augmentation essential for factual applications rather than optional.

Best-Fit Workloads

Where this model earns its place

01
Visual question answering and document analysis Text-plus-image input makes Gemma 3 27B suitable for VQA, image description, and document understanding, with a SigLIP-based vision encoder and MMMU around 64.9. Combined with the 128K context, it can analyze long documents that mix text and figures in a single request. It fits internal tools that extract or summarize from scanned pages, screenshots, and mixed-media inputs where self-hosting is preferred.
02
Multilingual generation and translation-style tasks With support for over 140 languages and strong post-training on multilingual data, Gemma 3 27B handles generation, summarization, and cross-lingual tasks across a wide language range. Global-MMLU-Lite of 75.1 supports broad multilingual competence. It suits products serving international audiences that need on-prem or license-permissive multilingual inference rather than a proprietary API.
03
Retrieval-augmented generation (RAG) A 74.9 FACTS Grounding score alongside a low SimpleQA indicates the model performs best when answers are grounded on provided context. The 128K window lets you feed substantial retrieved passages per query. This makes Gemma 3 27B a practical open-weight generator inside RAG pipelines, where retrieval compensates for the August 2024 cutoff and weak unaided recall.
04
Self-hosted, cost-controlled inference Because Gemma 3 27B is the most capable Gemma 3 size that runs on a single GPU — including consumer cards with quantization-aware training — it fits teams that need predictable, self-managed inference. Open weights under the Gemma license permit commercial use and full deployment control, avoiding per-token API dependence for high-volume or privacy-sensitive workloads.
Pricing in anytokens via AnyAPI
Input
0.48
₳
Output
0.96
₳
Cache write
—
₳
Cache read
—
₳

Integration

Access Gemma 3 27B Instruct via AnyAPI.ai

Access Gemma 3 27B Instruct through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate Gemma 3 27B Instruct without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access Gemma 3 27B Instruct and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test Gemma 3 27B Instruct against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use Gemma 3 27B Instruct from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use Gemma 3 27B Instruct for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

Google specifies a 128K-token context window for the 4B, 12B, and 27B Gemma 3 models. Some hosted API providers serve a larger 262,144-token context. Either way, this covers long documents, extended conversations, and substantial retrieved context in a single request.

Yes. Gemma 3 27B is multimodal for input, accepting both text and images and generating text output. A SigLIP vision encoder enables visual question answering, image description, and document analysis. It does not generate images — output is text only.

No. Gemma 3 27B is a non-reasoning model that produces direct responses without an extended chain-of-thought or thinking mode. It performs well on math and knowledge for its size, but for workloads that depend on step-by-step deliberation or agentic tool use, a dedicated reasoning model is a better fit.

Gemma 3 27B is released with open weights under the Gemma license, which permits commercial use and self-hosting. The weights are publicly downloadable, so you can run it on your own infrastructure or access it through hosted API providers.

Google's model card states the training data knowledge cutoff is August 2024. The model has no built-in web search, so for current information you should pair it with a retrieval-augmented generation pipeline, which also improves factual reliability given its low unaided recall.

* Benchmark data source: Artificial Analysis artificialanalysis.ai