OpenAI
•
GPT-5.2 Codex (Xhigh)
•
Released 
December 2025

OpenAI
GPT-5.2 Codex (Xhigh)

OpenAI's coding-specialized model for long-horizon agentic software engineering, large refactors, migrations, and structured code review.

Modality:
Text
Image
PDF
model ID
openai/gpt-5.2-codex

Output Speed *

N/A
tok/s

Intelligence Index *

28.5
/ 100

Context Window *

400000
tokens

Input price

10.5
Anytoken

Output price

84
Anytoken
GPT-5.2-Codex: A Coding-Specialized Model Built for Long-Horizon Agentic Engineering GPT-5.2-Codex is OpenAI's coding-specialized model, a version of GPT-5.2 further optimized for agentic software engineering. It adds improvements on long-horizon work through context compaction, stronger performance on large code changes like refactors and migrations, improved performance in Windows environments, and significantly stronger cybersecurity capabilities. It sits below general-purpose GPT-5.2 for open-ended reasoning but ahead of it for sustained coding agents. It fits teams running multi-step engineering tasks, automated refactors, and code review inside Codex-style harnesses rather than general chat workloads. Access GPT-5.2-Codex through the AnyAPI.ai API

Performance

Where GPT-5.2-Codex Earns Its Place: Sustained Agentic Coding

GPT-5.2-Codex is tuned for extended, multi-step engineering rather than single-turn answers. It improves long-horizon work through context compaction and handles large code changes like refactors and migrations more reliably. Independently, at 158 tokens per second it is notably fast. For coding agents, sustained throughput over long runs matters more than a marginally higher intelligence score, because the model must plan, edit, and validate across many turns. In production, this translates to fewer stalled runs on project-scale tasks and more usable end-to-end diffs from a single invocation.

Benchmarks

GPT-5.2-Codex Benchmarks: Coding-Weighted, Not Generalist-Leading

On independent aggregation, <cite index="28-1">GPT-5.2-Codex scores 72.4 on SWE-bench Verified.</cite> On general reasoning, <cite index="21-1,21-2">it achieves a score of 40 on the Artificial Analysis Intelligence Index, a composite covering reasoning, knowledge, mathematics, and coding.</cite> Artificial Analysis positions it <cite index="21-14">among the leading models in intelligence, but somewhat expensive relative to others of similar price.</cite> Read together, the picture is a coding-weighted specialist: strong on real-world software engineering tasks, competitive but not category-leading on broad generalist reasoning. Choose it for engineering agents, not for maximizing open-domain reasoning scores.

Output Speed

*
N/A
tok/s

Intelligence Index

*
28.5
/ 100

MMLU *

Broad world knowledge and problem-solving
0
%

GPQA *

PhD-level scientific reasoning across physics, biology, chemistry.
90
%

HLE *

Adherence to multi-step structured instructions.
36
%

LiveCodeBench *

Tool-calling reliability in long agentic loops.
0
%

Technical Specifications

What the model supports

GPT-5.2-Codex accepts text and image input and returns text only; audio and video are not supported. <cite index="5-3">It has a 400,000-token context window with a maximum output of 128,000 tokens.</cite> The large context is the defining production trait: it lets an agent hold full repositories, logs, and prior turns in a single run, which is what enables the long-horizon refactor and migration behavior. Image input covers screenshots for UI work, but it is not a general multimodal model. Reasoning effort is configurable, letting teams trade latency for depth per task.
Verified Specifications — 
GPT-5.2 Codex (Xhigh)
*
Input modalities
Text
Image
PDF
output modalities
Text
Context window
400000
 tokens
Maximum output tokens
128000
Reasoning
Yes
Knowledge cutoff
December 2025
Pricing (standard)
10.5
 AnyTokens in
 / 
84
 AnyTokens out

Limitations & Trade-offs

Where GPT-5.2 Codex (Xhigh) falls short

1
High latency at high reasoning effort. Long-horizon runs at elevated reasoning settings sacrifice responsiveness; independent measurements of GPT-5.2 in xhigh mode show time-to-first-token far higher than fast general models. This is acceptable for autonomous background engineering agents but unsuitable for interactive, latency-sensitive tooling. For real-time completions or chat-speed assistance, a faster model or a lower reasoning effort is preferable.
2
Expensive long outputs. GPT-5.2-Codex sits in a higher output-cost tier than its predecessor GPT-5.1-Codex, and it can be verbose. Because coding agents generate large diffs and long reasoning traces, output-heavy workloads accumulate cost quickly. For high-volume generation where cost dominates, the older Codex model or a lower reasoning effort is more economical.
3
Narrow specialization. OpenAI recommends the Codex family only for agentic coding in Codex or Codex-like environments, not general-purpose use. Its Intelligence Index of 40 is competitive but not the highest among frontier reasoning models. For open-domain reasoning, knowledge work, or chat, a general-purpose GPT-5.x model is the better fit.
4
Text-and-image input only, text output. GPT-5.2-Codex accepts text and image (screenshots) but does not support audio or video, and outputs text only. Teams needing audio, document-native pipelines beyond images, or non-text generation must pair it with other models. The image support is oriented toward UI development rather than broad multimodal understanding.

Best-Fit Workloads

Where this model earns its place

01

Repository-scale refactors and migrations

‍
The 400,000-token context and long-horizon design target exactly this. GPT-5.2-Codex includes improvements on long-horizon work through context compaction and stronger performance on large code changes like refactors and migrations. An agent can hold broad codebase context and iterate across many turns. Best suited to autonomous or supervised background runs where completion quality matters more than immediate latency.

02

Structured code review

‍
The model is trained to reason over dependencies and validate behavior against tests, making it well-suited to automated pull-request review that flags real defects rather than surface issues. Its improved steerability and closer instruction adherence over GPT-5.1-Codex help enforce team-specific review standards. Useful in CI pipelines that gate merges, though verbose output should be constrained for concise review comments.

03

Long-running coding agents

‍
GPT-5.2-Codex is built to sustain extended, multi-hour engineering runs while adapting reasoning depth per step. Configurable reasoning.effort lets an orchestrator use lower effort for small edits and xhigh for hard subtasks. Fits Codex CLI, IDE extension, and cloud-task workflows where the model plans, edits, installs dependencies, and validates autonomously across a session.

04

Defensive cybersecurity engineering

‍
GPT-5.2-Codex has stronger cybersecurity capabilities than any prior OpenAI model.It is very capable in the cybersecurity domain but does not reach High capability under the Preparedness Framework. This supports vulnerability triage and defensive code review; the dual-use nature means deployment should follow the provider's safety guidance and appropriate access controls.

Pricing in anytokens via AnyAPI
Input
10.5
₳
Output
84
₳
Cache write
—
₳
Cache read
—
₳

Integration

Access GPT-5.2 Codex (Xhigh) via AnyAPI.ai

Access GPT-5.2 Codex (Xhigh) through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate GPT-5.2 Codex (Xhigh) without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access GPT-5.2 Codex (Xhigh) and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test GPT-5.2 Codex (Xhigh) against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use GPT-5.2 Codex (Xhigh) from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use GPT-5.2 Codex (Xhigh) for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

GPT-5.2-Codex is a coding-specialized model for agentic software engineering. It targets long-horizon tasks such as repository-scale refactors, migrations, feature development, and structured code review inside Codex or Codex-like environments. OpenAI recommends the Codex family only for agentic coding, not general-purpose chat or open-domain reasoning.

GPT-5.2-Codex has a 400,000-token context window and a maximum output of 128,000 tokens. The large context lets an agent hold broad codebase context, logs, and prior turns in a single run, which enables its long-horizon refactor and migration behavior.

Yes. GPT-5.2-Codex accepts text and image input and returns text output. Image input is oriented toward UI development, such as screenshots. Audio and video are not supported, and the model does not generate non-text output.

GPT-5.2-Codex is an upgraded version of GPT-5.1-Codex. It is more steerable, follows developer instructions more closely, and produces cleaner code, with improvements in long-horizon work, large refactors, Windows environments, and cybersecurity. It carries an August 2025 knowledge cutoff versus September 2024, but sits in a higher output-cost tier.

GPT-5.2-Codex supports configurable reasoning effort via the reasoning.effort parameter, with low, medium, high, and xhigh settings. Lower effort reduces latency and cost for small tasks; xhigh maximizes depth for large, long-horizon runs. Most published benchmark figures use the xhigh setting.

* Benchmark data source: Artificial Analysis artificialanalysis.ai