OpenAI
•
GPT-5 Image
•
Released 
August 2025

OpenAI
GPT-5 Image

GPT-5 reasoning fused with GPT Image 1 generation, for workflows that reason and produce images in one call.

Modality:
Text
Image
PDF
model ID
openai/gpt-5-image

Output Speed *

N/A
tok/s

Intelligence Index *

N/A
/ 100

Context Window *

400000
tokens

Input price

60
Anytoken

Output price

60
Anytoken
GPT-5 Image: Reason and Generate Images in a Single Multimodal Call GPT-5 Image is OpenAI's combined endpoint that pairs the GPT-5 language model with GPT Image 1's image generation. It accepts text, images, and files such as PDFs as input, and returns both images and text. Positioned as a multimodal variant of the GPT-5 line rather than a standalone image endpoint, it lets a single request reason over context and emit generated or edited imagery. It suits applications that need instruction-following, reliable in-image text rendering, and detailed image editing without splitting work across a separate LLM and image API. Start building with GPT-5 Image through AnyAPI.ai

Performance

Combining GPT-5 Reasoning with GPT Image 1 Rendering

GPT-5 Image's advantage is producing images inside a reasoning flow rather than as an isolated call. It carries GPT-5's improvements in reasoning and instruction following while incorporating GPT Image 1's text rendering and detailed image editing. Because prompt interpretation and generation happen in one model, applications that translate structured instructions into precise visuals — including reliable in-image text — avoid the drift that occurs when an LLM's plan is passed to a separate image model. The practical consequence is fewer round-trips for design, annotation, and edit-in-context tasks, at the cost of higher per-request overhead than a pure image endpoint.

Benchmarks

Where GPT-5 Image's Underlying Model Stands

OpenRouter reports independent benchmarks from Artificial Analysis for GPT-5 Image. On the underlying GPT-5 model, Artificial Analysis measured a strong Intelligence Index and notably high long-context reasoning, where GPT-5 occupied top positions on their AA-LCR evaluation. Independent image-generation research has also evaluated the GPT-5-Image family against competing systems such as Gemini and Seedream. These figures reflect the language and generation capabilities the combined endpoint inherits; treat them as directional for image-plus-reasoning workloads rather than a single unified score, since text and image quality are measured separately.

Output Speed

*
N/A
tok/s

Intelligence Index

*
N/A
/ 100

MMLU *

Broad world knowledge and problem-solving
0
%

GPQA *

PhD-level scientific reasoning across physics, biology, chemistry.
0
%

HLE *

Adherence to multi-step structured instructions.
0
%

LiveCodeBench *

Tool-calling reliability in long agentic loops.
0
%

Technical Specifications

What the model supports

GPT-5 Image exposes a 400,000-token context window with up to 128,000 output tokens. Unlike text-only GPT-5 variants, this endpoint returns images and text, not text alone. The most production-relevant distinction is modality: it accepts text, images, and PDFs and emits generated or edited images, making it a genuine image-output model rather than a vision-input-only one. Structured outputs via JSON schema are supported. Note that function-calling tools are not necessarily available on the combined image endpoint, so agentic tool loops may require a separate mainline GPT-5 model.
Verified Specifications — 
GPT-5 Image
*
Input modalities
Text
Image
PDF
output modalities
Text
Image
Context window
400000
 tokens
Maximum output tokens
128000
Reasoning
Yes
Knowledge cutoff
August 2025
Pricing (standard)
60
 AnyTokens in
 / 
60
 AnyTokens out

Limitations & Trade-offs

Where GPT-5 Image falls short

1
Higher per-request overhead than a pure image endpoint. GPT-5 Image runs a full GPT-5 reasoning pass around generation, so every request carries language-model cost and latency even for simple text-to-image tasks. OpenAI's own guidance suggests calling a GPT Image endpoint directly for image-only workloads to avoid this LLM overhead. Reserve GPT-5 Image for cases where the reasoning genuinely adds value; high-volume, straightforward generation is better served elsewhere.
2
Function calling may be unavailable on the combined endpoint. On the sibling GPT-5.4 Image 2 endpoint, tools are not accepted, meaning function calling is unavailable there. If your architecture depends on agentic tool loops interleaved with image generation, you likely need to orchestrate a separate mainline GPT-5 model for tool use and delegate generation to the image endpoint, rather than expecting one call to do both.
3
Image generation, not native image editing controls of a dedicated pipeline. While GPT Image 1 supports detailed editing, the combined endpoint's editing features are more constrained than the newer Image 2 generation, which adds multi-reference support. Teams doing heavy compositing or reference-driven editing may find the newer sibling or a dedicated image pipeline more capable.
4
Older reasoning and generation cores. GPT-5 Image is built on the original GPT-5 (knowledge cutoff September 2024) and GPT Image 1. Newer GPT-5-series models and GPT Image 2 offer stronger reasoning and image quality. For applications sensitive to recent knowledge or top-tier generation fidelity, a more recent combined or standalone model is preferable.

Best-Fit Workloads

Where this model earns its place

01

Reason-then-generate design assistants

‍
Applications where a user describes intent in natural language and expects a generated visual benefit from doing interpretation and generation in one model. GPT-5 Image applies GPT-5's instruction following before emitting an image, reducing the semantic drift seen when a separate LLM plan is handed to an isolated image model. Suited to creative tools and marketing-asset generators that need prompts understood in context.

02

In-image text and typographic assets

‍
GPT Image 1's reliable text rendering makes GPT-5 Image appropriate for generating images that must contain legible, correct text — posters, social cards, labeled diagrams, and mockups. Combined with reasoning, the model can compose layout instructions and produce imagery with intended wording, a task where many image models fail. Validate output text for accuracy in production, as no generator is perfectly reliable.

03

Document-grounded visual output

‍
Because the endpoint accepts PDFs and images as input within a 400K context, it can reason over a supplied document and produce an accompanying visual — for example summarizing a report and generating an illustrative graphic or annotated edit. The large context window is the differentiator here, letting long source material inform generation in a single request.

04

Contextual image editing

‍
GPT-5 Image can take an input image plus instructions and return an edited result, using GPT Image 1's detailed editing. This fits workflows like iterative revision of a supplied asset where the model must understand both the image and a textual edit request. For heavy multi-reference compositing, the newer GPT-5.4 Image 2 with sixteen-reference support is the stronger option.

Pricing in anytokens via AnyAPI
Input
60
₳
Output
60
₳
Cache write
—
₳
Cache read
—
₳

Integration

Access GPT-5 Image via AnyAPI.ai

Access GPT-5 Image through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate GPT-5 Image without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access GPT-5 Image and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test GPT-5 Image against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use GPT-5 Image from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use GPT-5 Image for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

GPT-5 Image combines OpenAI's GPT-5 language model with GPT Image 1's generation capabilities in a single endpoint that returns images and text. GPT Image 1 alone is a dedicated image model. GPT-5 Image adds full reasoning around generation, so it interprets complex instructions before producing imagery, at the cost of higher per-request overhead.

GPT-5 Image provides a 400,000-token context window and supports up to 128,000 output tokens per request. The context window is shared across text, image, and PDF inputs, making it suitable for reasoning over long documents while generating an accompanying image.

GPT-5 Image accepts text, images, and files such as PDFs as input, and returns both images and text as output. This makes it an image-output model, unlike text-only GPT-5 variants that accept images as input but only return text.

GPT-5 Image supports structured outputs via a JSON schema. Function calling may not be available on the combined image endpoint, mirroring the sibling GPT-5.4 Image 2 endpoint where tools are not accepted. For agentic tool loops, orchestrate a separate mainline GPT-5 model alongside image generation.

Use GPT-5 Image when generation must be tied to interpreted context — for example following complex instructions, rendering correct in-image text, or producing a visual grounded in a supplied document. For simple, high-volume text-to-image with no reasoning, OpenAI recommends a direct GPT Image endpoint to avoid the language-model overhead.

* Benchmark data source: Artificial Analysis artificialanalysis.ai