OpenAI
•
gpt-oss-120b (High)
•
Released 
August 2025

OpenAI
gpt-oss-120b (High)

OpenAI's open-weight MoE reasoning model with configurable effort, near-o4-mini accuracy, and a permissive Apache 2.0 license.

Modality:
Text
PDF
model ID
openai/gpt-oss-120b

Output Speed *

199.82
tok/s

Intelligence Index *

11.6
/ 100

Context Window *

131072
tokens

Input price

0.234
Anytoken

Output price

1.14
Anytoken
gpt-oss-120b: Open-Weight Reasoning With Configurable Effort and Full Chain-of-Thought gpt-oss-120b is OpenAI's larger open-weight reasoning model, released under Apache 2.0. It is a Mixture-of-Experts transformer with 117B total parameters and 5.1B active per token, designed to fit on a single 80GB GPU. Within the gpt-oss family it sits above the lighter gpt-oss-20b. Its distinguishing features are configurable reasoning effort (low, medium, high), full chain-of-thought access, and native agentic tool use, reaching near-parity with OpenAI o4-mini on core reasoning benchmarks. It suits self-hosted, cost-controlled reasoning and agentic workloads where weight ownership and transparency matter. Start building with the gpt-oss-120b API on AnyAPI.ai

Performance

Reasoning Accuracy Near o4-mini From an Open-Weight Model

gpt-oss-120b is built for reasoning and agentic tool use, not just chat. OpenAI reports it achieves near-parity with o4-mini on core reasoning benchmarks while running on a single 80GB GPU, and independent testing places it at 24 on the Artificial Analysis Intelligence Index versus a ~9 median for similarly sized open-weight models. Configurable reasoning effort (low, medium, high) lets teams trade accuracy against latency and token spend. The production consequence: strong math, coding, and agentic reasoning are available as downloadable weights, enabling self-hosting and fine-tuning without depending on a proprietary endpoint.

Benchmarks

How gpt-oss-120b Scores Across Reasoning, Coding and Agentic Evaluations

On OpenAI's published evaluations, gpt-oss-120b at high effort reaches roughly 90% MMLU, 62.4% on SWE-Bench Verified, and 95.8% on AIME 2024 (no tools), rising to 96.6% with tools. Accuracy scales sharply with reasoning effort—for example, AIME 2024 climbs from 56.3% at low to 95.8% at high—so effort choice materially affects output quality. Independent Artificial Analysis testing rates it notably fast (~180 tokens/sec median) but somewhat verbose, generating well above the median token count on its Intelligence Index suite, which increases output-token consumption on reasoning tasks.

Output Speed

*
199.82
tok/s

Intelligence Index

*
11.6
/ 100

MMLU *

Broad world knowledge and problem-solving
81
%

GPQA *

PhD-level scientific reasoning across physics, biology, chemistry.
78
%

HLE *

Adherence to multi-step structured instructions.
20
%

LiveCodeBench *

Tool-calling reliability in long agentic loops.
88
%

Technical Specifications

What the model supports

gpt-oss-120b is a text-only model with a 131,072-token context window and matching maximum output, based on a Mixture-of-Experts architecture (117B total, 5.1B active). It provides configurable reasoning effort, full chain-of-thought, native function calling, web browsing, Python execution, and structured outputs. The two most production-relevant characteristics are the open-weight Apache 2.0 release, which permits commercial self-hosting and fine-tuning, and the configurable reasoning effort, which directly controls latency and output-token cost. It uses OpenAI's harmony response format and is compatible with the Responses API.
Verified Specifications — 
gpt-oss-120b (High)
*
Input modalities
Text
PDF
output modalities
Text
Context window
131072
 tokens
Maximum output tokens
128000
Reasoning
Yes
Knowledge cutoff
August 2025
Pricing (standard)
0.234
 AnyTokens in
 / 
1.14
 AnyTokens out

Quickstart

Sample code for gpt-oss-120b (High)

import requests

url = "https://api.anyapi.ai/v1/chat/completions"

payload = {
    "stream": False,
    "tool_choice": "auto",
    "logprobs": False,
    "model": "gpt-oss-120b",
    "messages": [
        {
            "role": "user",
            "content": "Hello"
        }
    ]
}
headers = {
    "Authorization": "Bearer AnyAPI_API_KEY",
    "Content-Type": "application/json"
}

response = requests.post(url, json=payload, headers=headers)

print(response.json())
import requests url = "https://api.anyapi.ai/v1/chat/completions" payload = { "stream": False, "tool_choice": "auto", "logprobs": False, "model": "gpt-oss-120b", "messages": [ { "role": "user", "content": "Hello" } ] } headers = { "Authorization": "Bearer AnyAPI_API_KEY", "Content-Type": "application/json" } response = requests.post(url, json=payload, headers=headers) print(response.json())
View docs
Copy
Code is copied
const url = 'https://api.anyapi.ai/v1/chat/completions';
const options = {
  method: 'POST',
  headers: {Authorization: 'Bearer AnyAPI_API_KEY', 'Content-Type': 'application/json'},
  body: '{"stream":false,"tool_choice":"auto","logprobs":false,"model":"gpt-oss-120b","messages":[{"role":"user","content":"Hello"}]}'
};

try {
  const response = await fetch(url, options);
  const data = await response.json();
  console.log(data);
} catch (error) {
  console.error(error);
}
const url = 'https://api.anyapi.ai/v1/chat/completions'; const options = { method: 'POST', headers: {Authorization: 'Bearer AnyAPI_API_KEY', 'Content-Type': 'application/json'}, body: '{"stream":false,"tool_choice":"auto","logprobs":false,"model":"gpt-oss-120b","messages":[{"role":"user","content":"Hello"}]}' }; try { const response = await fetch(url, options); const data = await response.json(); console.log(data); } catch (error) { console.error(error); }
View docs
Copy
Code is copied
curl --request POST \
  --url https://api.anyapi.ai/v1/chat/completions \
  --header 'Authorization: Bearer AnyAPI_API_KEY' \
  --header 'Content-Type: application/json' \
  --data '{
  "stream": false,
  "tool_choice": "auto",
  "logprobs": false,
  "model": "gpt-oss-120b",
  "messages": [
    {
      "role": "user",
      "content": "Hello"
    }
  ]
}'
curl --request POST \ --url https://api.anyapi.ai/v1/chat/completions \ --header 'Authorization: Bearer AnyAPI_API_KEY' \ --header 'Content-Type: application/json' \ --data '{ "stream": false, "tool_choice": "auto", "logprobs": false, "model": "gpt-oss-120b", "messages": [ { "role": "user", "content": "Hello" } ] }'
View docs
Copy
Code is copied
View docs
Code examples coming soon...

Limitations & Trade-offs

Where gpt-oss-120b (High) falls short

1
Verbose reasoning outputs. Independent Artificial Analysis testing found gpt-oss-120b generates well above the median token count on its Intelligence Index (roughly 88M vs a ~57M median) and describes it as somewhat verbose. Because output tokens dominate reasoning cost, this raises spend and latency on long chains of thought. Workloads with tight per-request budgets or strict latency targets should cap reasoning effort or prefer a leaner model.
2
Text-only, no multimodal input. gpt-oss-120b accepts and produces text only—there is no image, audio, or document-vision capability. Applications requiring screenshot understanding, chart reading, or visual agents must pair it with a separate vision model or choose a multimodal alternative. This is a hard limitation of the released weights, not a configuration option.
3
Self-hosting and safety responsibility. As an open-weight release, gpt-oss-120b shifts operational and safety responsibility to the deployer. OpenAI notes open models present a different risk profile: once released, they can be fine-tuned to bypass safeguards, and there is no server-side ability to revoke access or add mitigations. Teams must implement their own system-level protections, which adds engineering overhead versus a managed proprietary endpoint.
4
Performance depends heavily on provider and effort. Independent measurements show output speed for gpt-oss-120b varies enormously across hosts—from tens of tokens per second on some GPU clouds to well over a thousand on specialized hardware—and accuracy scales sharply with reasoning effort. There is no single canonical latency or quality figure; results depend on the chosen provider and effort setting, complicating benchmarking and capacity planning.

Best-Fit Workloads

Where this model earns its place

01
Agentic tool-use pipelines gpt-oss-120b was explicitly optimized for agentic capabilities including function calling, browsing, and Python execution, and matches or exceeds o4-mini on tool-calling evaluations like TauBench. Full chain-of-thought access aids debugging of multi-step agent loops. It fits autonomous research assistants and tool-orchestrating agents. Note that verbose reasoning increases token consumption per step, so cap effort where tasks are simple.
02
Self-hosted reasoning and math With near-o4-mini reasoning accuracy and strong AIME and GPQA Diamond results at high effort, gpt-oss-120b suits math-heavy and STEM reasoning tasks that teams need to run on their own infrastructure. The Apache 2.0 license permits commercial self-hosting on a single 80GB GPU. This is a fit for regulated or privacy-sensitive environments requiring full data control rather than a proprietary API.
03
Software engineering assistance gpt-oss-120b scores 62.4% on SWE-Bench Verified at high effort and matches or exceeds o4-mini on competition coding, making it viable for code generation, bug fixing, and repository-level engineering tasks. Full chain-of-thought helps inspect its problem-solving. Independent leaderboards rate pure coding at scale below leading proprietary and specialist models, so pair with human review for complex changes.
04
Fine-tuned domain models Because the weights are open and fully fine-tunable under Apache 2.0, gpt-oss-120b is a strong base for building customized domain models—for example specialized reasoning, health, or internal-tooling assistants. Teams gain parameter-level customization impossible with closed APIs. This requires ML infrastructure and safety tuning, so it fits organizations with the engineering capacity to own the training and deployment stack.
Pricing in anytokens via AnyAPI
Input
0.234
₳
Output
1.14
₳
Cache write
—
₳
Cache read
—
₳

Integration

Access gpt-oss-120b (High) via AnyAPI.ai

Access gpt-oss-120b (High) through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate gpt-oss-120b (High) without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access gpt-oss-120b (High) and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test gpt-oss-120b (High) against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use gpt-oss-120b (High) from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use gpt-oss-120b (High) for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

gpt-oss-120b has a 131,072-token (128K) context window, using RoPE positional encoding. Managed API listings report a matching maximum output of 131,072 tokens. This applies to a single request covering both prompt and completion, so long reasoning outputs consume the same budget as input context.

gpt-oss-120b is released as an open-weight model under the Apache 2.0 license, which permits commercial use, self-hosting, and fine-tuning. The weights are freely downloadable and ship natively quantized to fit a single 80GB GPU. OpenAI also open-sourced the harmony format and o200k_harmony tokenizer.

Yes. gpt-oss-120b supports native function calling, web browsing, Python code execution, and structured outputs. It offers configurable reasoning effort at low, medium, and high levels, plus full chain-of-thought access. Accuracy scales notably with effort—benchmarks like AIME rise substantially from low to high effort.

OpenAI reports gpt-oss-120b achieves near-parity with o4-mini on core reasoning benchmarks, matching or exceeding it on competition coding, MMLU, HLE, and tool calling, while running on a single 80GB GPU. The key difference is that gpt-oss-120b is open-weight and self-hostable, whereas o4-mini is proprietary.

No. gpt-oss-120b is a text-only model for both input and output. It does not accept images, audio, video, or document-vision inputs. Multimodal applications must pair it with a separate vision or audio model, or select a natively multimodal alternative.

* Benchmark data source: Artificial Analysis artificialanalysis.ai