OpenAI
•
GPT-5 (High)
•
Released 
August 2025

OpenAI
GPT-5 (High)

OpenAI's reasoning-first flagship with configurable effort, 400K context, and strong real-world coding and agentic tool use.

Modality:
Text
Image
PDF
model ID
openai/gpt-5

Output Speed *

0.00
tok/s

Intelligence Index *

23
/ 100

Context Window *

400000
tokens

Input price

7.5
Anytoken

Output price

60
Anytoken
GPT-5: Configurable Reasoning for Coding and Long-Horizon Agents GPT-5 is OpenAI's reasoning-first flagship, released August 2025 for coding, complex reasoning, and agentic task execution. It pairs a 400,000-token context window with configurable reasoning effort (minimal to high) and a verbosity control, letting one model span quick responses and deep multi-step problem solving. Its strongest evidence is real-world coding and tool orchestration: it scored 74.9% on SWE-bench Verified and 88% on Aider Polyglot with reasoning enabled. GPT-5 suits teams building coding agents and long chains of tool calls. OpenAI now positions it as its previous flagship, recommending later GPT-5.x models for new work. Start building with the GPT-5 API on AnyAPI.ai

Performance

Where GPT-5 Earns Its Place: Coding and Tool Orchestration

GPT-5 is built for real-world software work and long chains of tool calls rather than single-shot chat. With reasoning enabled it reached 74.9% on SWE-bench Verified and 88% on Aider Polyglot, both measuring code changes against realistic repositories. That combination matters because production coding agents fail on multi-step edits and stale context, not toy prompts; a model that plans, calls tools, and revises across a 400K-token window resolves more issues end-to-end with less human intervention. The practical consequence is fewer supervised retries in agent loops, though enabling reasoning increases latency and token consumption on each request.

Benchmarks

GPT-5 Independent Benchmarks: Intelligence, Speed and Latency

Artificial Analysis rates GPT-5 (high) as above average in intelligence and faster than average, while noting it is somewhat verbose. On OpenAI's first-party API at high reasoning effort, independent measurement recorded roughly 95 output tokens per second and a time to first token near 91 seconds, reflecting the thinking phase reasoning models spend before answering. The verbosity and long time-to-first-token are the key selection signals: GPT-5 favors thoroughness over snappy responses, so it fits asynchronous agent and batch workloads better than latency-critical, interactive ones. Lowering reasoning effort trades intelligence for faster first tokens.

Output Speed

*
0.00
tok/s

Intelligence Index

*
23
/ 100

MMLU *

Broad world knowledge and problem-solving
87
%

GPQA *

PhD-level scientific reasoning across physics, biology, chemistry.
85
%

HLE *

Adherence to multi-step structured instructions.
28
%

LiveCodeBench *

Tool-calling reliability in long agentic loops.
85
%

Technical Specifications

What the model supports

GPT-5 accepts text and image input and returns text only; it is not a native audio, video, or image-generation model. It provides a 400,000-token context window and can produce up to 128,000 tokens in a single response — a large output ceiling that is separate from and much smaller than total context. The most production-relevant control is reasoning effort (minimal, low, medium, high) plus a verbosity parameter, letting teams tune the intelligence-versus-latency trade-off per request. Knowledge is current to September 30, 2024, so recent facts require retrieval or tool augmentation.
Verified Specifications — 
GPT-5 (High)
*
Input modalities
Text
Image
PDF
output modalities
Text
Context window
400000
 tokens
Maximum output tokens
128000
Reasoning
Yes
Knowledge cutoff
August 2025
Pricing (standard)
7.5
 AnyTokens in
 / 
60
 AnyTokens out

Quickstart

Sample code for GPT-5 (High)

‍

import requests

url = "https://api.anyapi.ai/v1/chat/completions"

payload = {
    "model": "gpt-5",
    "messages": [
        {
            "role": "user",
            "content": [
                {
                    "type": "text",
                    "text": "Text prompt"
                },
                {
                    "image_url": { "url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg" },
                    "type": "image_url"
                }
            ]
        }
    ]
}
headers = {
    "Authorization": "Bearer  AnyAPI_API_KEY",
    "Content-Type": "application/json"
}

response = requests.post(url, json=payload, headers=headers)

print(response.json())
import requests url = "https://api.anyapi.ai/v1/chat/completions" payload = { "model": "gpt-5", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Text prompt" }, { "image_url": { "url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg" }, "type": "image_url" } ] } ] } headers = { "Authorization": "Bearer AnyAPI_API_KEY", "Content-Type": "application/json" } response = requests.post(url, json=payload, headers=headers) print(response.json())
View docs
Copy
Code is copied
const url = 'https://api.anyapi.ai/v1/chat/completions';
const options = {
  method: 'POST',
  headers: {Authorization: 'Bearer  AnyAPI_API_KEY', 'Content-Type': 'application/json'},
  body: '{"model":"gpt-5","messages":[{"role":"user","content":[{"type":"text","text":"Text prompt"},{"image_url":{"url":"https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"},"type":"image_url"}]}]}'
};

try {
  const response = await fetch(url, options);
  const data = await response.json();
  console.log(data);
} catch (error) {
  console.error(error);
}
const url = 'https://api.anyapi.ai/v1/chat/completions'; const options = { method: 'POST', headers: {Authorization: 'Bearer AnyAPI_API_KEY', 'Content-Type': 'application/json'}, body: '{"model":"gpt-5","messages":[{"role":"user","content":[{"type":"text","text":"Text prompt"},{"image_url":{"url":"https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"},"type":"image_url"}]}]}' }; try { const response = await fetch(url, options); const data = await response.json(); console.log(data); } catch (error) { console.error(error); }
View docs
Copy
Code is copied
curl --request POST \
  --url https://api.anyapi.ai/v1/chat/completions \
  --header 'Authorization: Bearer  AnyAPI_API_KEY' \
  --header 'Content-Type: application/json' \
  --data '{
  "model": "gpt-5",
  "messages": [
    {
      "role": "user",
      "content": [
        {
          "type": "text",
          "text": "Text prompt"
        },
        {
          "image_url": {
            "url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg"
          },
          "type": "image_url"
        }
      ]
    }
  ]
}'
curl --request POST \ --url https://api.anyapi.ai/v1/chat/completions \ --header 'Authorization: Bearer AnyAPI_API_KEY' \ --header 'Content-Type: application/json' \ --data '{ "model": "gpt-5", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Text prompt" }, { "image_url": { "url": "https://upload.wikimedia.org/wikipedia/commons/thumb/d/dd/Gfp-wisconsin-madison-the-nature-boardwalk.jpg/2560px-Gfp-wisconsin-madison-the-nature-boardwalk.jpg" }, "type": "image_url" } ] } ] }'
View docs
Copy
Code is copied
View docs
Code examples coming soon...

Limitations & Trade-offs

Where GPT-5 (High) falls short

1
High time to first token. GPT-5 is a reasoning model, and independent measurement of the first-party API at high effort recorded a time to first token near 91 seconds because the model reasons before answering. For interactive, latency-sensitive UX such as live chat or autocomplete, this is disqualifying at high effort; lower reasoning effort helps, but latency-critical products are usually better served by a faster non-reasoning model or a lighter tier.
2
Verbosity and output cost. Artificial Analysis characterizes GPT-5 as somewhat verbose, and reasoning tokens plus long completions inflate output volume. Because output is the more expensive side of usage, high-effort runs and long generations become comparatively costly at scale. High-volume, cost-sensitive workloads should cap verbosity, lower reasoning effort, or route simpler requests to a cheaper model.
3
Knowledge cutoff of September 30, 2024. GPT-5 has no built-in awareness of events, libraries, or API changes after that date. For anything current — recent framework versions, news, or evolving documentation — you must supply context through retrieval or tools, or the model will answer from stale training data. Newer family members carry later cutoffs.
4
Superseded flagship status. OpenAI now labels GPT-5 as its previous model and recommends later GPT-5.x releases for new work, which score higher on current independent coding-agent benchmarks. This does not break existing deployments, but teams beginning new projects should weigh whether building on a superseded flagship is the right long-term foundation versus adopting the current recommendation.

Best-Fit Workloads

Where this model earns its place

01

Autonomous coding agents

‍
GPT-5's 74.9% on SWE-bench Verified and 88% on Aider Polyglot with reasoning enabled make it a strong engine for agents that resolve GitHub issues, refactor across files, and ship patches end-to-end. Its support for long chains of tool calls lets an agent plan, run commands, inspect results, and revise. The 400K context holds substantial repository state. Budget for higher latency and token use per task at elevated reasoning effort.

02

Long-horizon tool-using assistants

‍
Because GPT-5 executes extended sequences of tool calls and follows multi-step plans, it suits back-office assistants that query systems, transform data, and produce structured results across many steps. Structured outputs and function calling keep responses machine-parseable inside pipelines. This fits asynchronous or queued work where a first-token delay of tens of seconds is acceptable, rather than real-time chat.

03

Large-context analysis and review

‍
The 400,000-token window supports codebase analysis, long technical documents, and multi-file review in a single request, and image input allows reasoning over diagrams and screenshots alongside text. Higher reasoning effort improves accuracy on dense material. Note the September 2024 knowledge cutoff means any current facts must come from the supplied context, not the model's memory.

04

Complex reasoning and STEM problem solving

‍
GPT-5 scored 94.6% on AIME 2025 without tools and 84.2% on MMMU for multimodal understanding, indicating strong math and structured-reasoning ability. This fits technical analysis, quantitative problem solving, and evaluation tasks where correctness outweighs speed. Enable higher reasoning effort for the hardest problems and accept the added latency and token cost that thorough reasoning requires.

Pricing in anytokens via AnyAPI
Input
7.5
₳
Output
60
₳
Cache write
—
₳
Cache read
—
₳

Integration

Access GPT-5 (High) via AnyAPI.ai

Access GPT-5 (High) through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate GPT-5 (High) without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access GPT-5 (High) and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test GPT-5 (High) against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use GPT-5 (High) from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use GPT-5 (High) for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

GPT-5 supports a 400,000-token context window, composed of up to 272,000 input tokens plus up to 128,000 output tokens. The 128,000-token maximum output is the most it can generate in a single response and is separate from total context. This is large enough for full codebase analysis, long documents, and multi-file review in one request.

Yes. GPT-5 is a reasoning model with a configurable reasoning-effort parameter offering minimal, low, medium, and high settings, plus a separate verbosity parameter. Higher effort improves accuracy on complex coding and math tasks but increases latency and token usage, letting developers tune the intelligence-versus-speed trade-off per request.

GPT-5 is designed for coding and agentic tasks. With reasoning enabled it scored 74.9% on SWE-bench Verified and 88% on Aider Polyglot, both measuring real code changes, and it supports long chains of tool calls for agent workflows. Note that OpenAI now recommends later GPT-5.x models for new coding-agent projects.

GPT-5 accepts text and image input and produces text output only. It does not natively generate images, audio, or video. Image input allows it to reason over diagrams, screenshots, and visual documents alongside text, which supports multimodal understanding tasks like MMMU where it scored 84.2%.

GPT-5's knowledge cutoff is September 30, 2024. It has no built-in awareness of events, library versions, or documentation changes after that date. For current information you must supply context via retrieval or tools. Later models in the GPT-5 family carry more recent cutoffs.

* Benchmark data source: Artificial Analysis artificialanalysis.ai