OpenAI
•
gpt-oss-20b (High)
•
Released 
August 2025

OpenAI
gpt-oss-20b (High)

OpenAI's compact open-weight reasoning model that runs on 16 GB hardware with configurable reasoning and native tool calling.

Modality:
Text
PDF
model ID
openai/gpt-oss-20b

Output Speed *

217.12
tok/s

Intelligence Index *

9
/ 100

Context Window *

131072
tokens

Input price

0.18
Anytoken

Output price

0.66
Anytoken
gpt-oss-20B: o3-mini-Class Reasoning You Can Run Almost Anywhere gpt-oss-20B is an open-weight reasoning model from OpenAI, released under the Apache 2.0 license. It is the smaller of the two gpt-oss models, built on a Mixture-of-Experts transformer with roughly 21B total parameters and only 3.6B active per token. That sparsity lets it run within 16 GB of memory while delivering results comparable to OpenAI o3-mini on common benchmarks. What distinguishes it is configurable reasoning effort, full chain-of-thought visibility, and native agentic tooling. It fits low-latency, cost-sensitive, and locally deployed reasoning and tool-use workloads. Integrate gpt-oss-20B via the AnyAPI.ai API in minutes.

Performance

Reasoning Density in a 16 GB Footprint

gpt-oss-20B's strength is reasoning quality relative to its size and deployment cost. OpenAI reports it matches or exceeds o3-mini on common evaluations, including competition mathematics and health, despite activating only 3.6B parameters per token. Independent testing places it above the median for comparable open-weight models on the Artificial Analysis Intelligence Index at high reasoning effort. This matters because it delivers o-series-style chain-of-thought locally, on consumer GPUs or laptops with roughly 16 GB of memory. In production, teams can run capable reasoning and tool-use pipelines without frontier-scale infrastructure, though larger models still lead on hard graduate-level tasks.

Benchmarks

Where gpt-oss-20B Lands in Independent Testing

On the Artificial Analysis Intelligence Index v4.1.1, gpt-oss-20B at high reasoning effort scores 15, above the comparable-model median of 9, while generating roughly 124 tokens per second. Independent evaluation notes the model is somewhat verbose, using more output tokens than the median to complete Intelligence Index tasks, which affects effective cost and latency. Community and provider benchmarks report MMLU around 85% and GPQA Diamond around 71.5%, where its smaller size shows against larger models. On competition mathematics (AIME) it performs strongly for its class, consistent with OpenAI's o3-mini-parity positioning.

Output Speed

*
217.12
tok/s

Intelligence Index

*
9
/ 100

MMLU *

Broad world knowledge and problem-solving
75
%

GPQA *

PhD-level scientific reasoning across physics, biology, chemistry.
69
%

HLE *

Adherence to multi-step structured instructions.
11
%

LiveCodeBench *

Tool-calling reliability in long agentic loops.
78
%

Technical Specifications

What the model supports

gpt-oss-20B is a text-only model with a 131,072-token context window and up to 32,768 output tokens on managed APIs. It supports configurable reasoning effort (low, medium, high), full chain-of-thought, native function calling, and Structured Outputs. The reasoning-effort control is the most production-relevant feature: low effort maximizes speed for routine dialogue, while high effort trades latency for accuracy on hard tasks. Note the maximum output ceiling is markedly lower than the full context window, so very long single generations must be planned around that 32K limit rather than the 131K input capacity.
Verified Specifications — 
gpt-oss-20b (High)
*
Input modalities
Text
PDF
output modalities
Text
Context window
131072
 tokens
Maximum output tokens
128000
Reasoning
Yes
Knowledge cutoff
August 2025
Pricing (standard)
0.18
 AnyTokens in
 / 
0.66
 AnyTokens out

Quickstart

Sample code for gpt-oss-20b (High)

import requests

url = "https://api.anyapi.ai/v1/chat/completions"

payload = {
    "model": "gpt-oss-20b",
    "messages": [
        {
            "role": "user",
            "content": "Hello"
        }
    ]
}
headers = {
    "Authorization": "Bearer AnyAPI_API_KEY",
    "Content-Type": "application/json"
}

response = requests.post(url, json=payload, headers=headers)

print(response.json())
import requests url = "https://api.anyapi.ai/v1/chat/completions" payload = { "model": "gpt-oss-20b", "messages": [ { "role": "user", "content": "Hello" } ] } headers = { "Authorization": "Bearer AnyAPI_API_KEY", "Content-Type": "application/json" } response = requests.post(url, json=payload, headers=headers) print(response.json())
View docs
Copy
Code is copied
const url = 'https://api.anyapi.ai/v1/chat/completions';
const options = {
  method: 'POST',
  headers: {Authorization: 'Bearer AnyAPI_API_KEY', 'Content-Type': 'application/json'},
  body: '{"model":"gpt-oss-20b","messages":[{"role":"user","content":"Hello"}]}'
};

try {
  const response = await fetch(url, options);
  const data = await response.json();
  console.log(data);
} catch (error) {
  console.error(error);
}
const url = 'https://api.anyapi.ai/v1/chat/completions'; const options = { method: 'POST', headers: {Authorization: 'Bearer AnyAPI_API_KEY', 'Content-Type': 'application/json'}, body: '{"model":"gpt-oss-20b","messages":[{"role":"user","content":"Hello"}]}' }; try { const response = await fetch(url, options); const data = await response.json(); console.log(data); } catch (error) { console.error(error); }
View docs
Copy
Code is copied
curl --request POST \
  --url https://api.anyapi.ai/v1/chat/completions \
  --header 'Authorization: Bearer AnyAPI_API_KEY' \
  --header 'Content-Type: application/json' \
  --data '{
  "model": "gpt-oss-20b",
  "messages": [
    {
      "role": "user",
      "content": "Hello"
    }
  ]
}'
curl --request POST \ --url https://api.anyapi.ai/v1/chat/completions \ --header 'Authorization: Bearer AnyAPI_API_KEY' \ --header 'Content-Type: application/json' \ --data '{ "model": "gpt-oss-20b", "messages": [ { "role": "user", "content": "Hello" } ] }'
View docs
Copy
Code is copied
View docs
Code examples coming soon...

Limitations & Trade-offs

Where gpt-oss-20b (High) falls short

1
Output length ceiling. Despite a 131,072-token context window, gpt-oss-20B caps output at 32,768 tokens on managed APIs—far below its input capacity. Long-form generation such as full document drafting or large code files must be chunked or planned around that limit. If your workload requires very long single completions, a model with a higher output ceiling is a better fit.
2
Text-only, no multimodality. gpt-oss-20B accepts and produces text only; it has no image, audio, or document-image input. It was trained on a mostly English text dataset. Vision-dependent workloads—OCR, chart reading, screenshot-based agents, or document understanding from images—need a multimodal model instead. Do not assume vision from other OpenAI models.
3
Weaker on hard graduate-level tasks. Because of its smaller size, gpt-oss-20B lags larger models on demanding evaluations; independent figures place GPQA Diamond around 71.5% and MMLU around 85%, and OpenAI notes the 20B trails on GPQA. For frontier scientific reasoning or the hardest analytical work, gpt-oss-120B or a larger proprietary model performs better.
4
Verbosity and reasoning overhead. Independent testing found gpt-oss-20B somewhat verbose, using more output tokens than the median to complete Intelligence Index tasks, and high reasoning effort further increases chain-of-thought length. Since output tokens drive cost and latency, this raises effective spend on reasoning-heavy calls. For latency-critical or budget-tight paths, use lower reasoning effort or a less verbose model.

Best-Fit Workloads

Where this model earns its place

01

Local and edge reasoning agents

‍
gpt-oss-20B's MoE sparsity and MXFP4 quantization let it run within roughly 16 GB of memory, enabling on-device agents on laptops or single consumer GPUs. Combined with native function calling and full chain-of-thought, it powers offline or privacy-sensitive assistants without cloud round-trips. Plan around the 32K output cap and provider-dependent latency.

02

Cost-sensitive high-volume tool-calling

‍
With only 3.6B active parameters, gpt-oss-20B is inexpensive to serve at scale, making it a strong default for high-throughput agentic pipelines that issue many function calls. OpenAI positions the gpt-oss family for strong tool use, and it supports web browsing, Python execution, and developer-defined functions. Use low or medium reasoning effort to control per-call cost.

03

Fine-tuned domain reasoning

‍
The Apache 2.0 license and full fine-tunability let teams adapt gpt-oss-20B to specialized domains without copyleft or patent risk. Researchers already LoRA-tune it for long-context and reasoning tasks. This suits companies needing a customizable, self-hostable reasoning model with predictable licensing, where a proprietary API would be restrictive or costly.

04

Structured extraction and workflow automation

‍
Native Structured Outputs and JSON-schema support make gpt-oss-20B suitable for turning unstructured text into typed records within automation pipelines. Its o3-mini-class reasoning handles multi-field extraction and conditional logic, while configurable reasoning effort keeps routine extractions fast. Keep individual outputs under the 32K-token ceiling for large batch jobs.

Pricing in anytokens via AnyAPI
Input
0.18
₳
Output
0.66
₳
Cache write
—
₳
Cache read
—
₳

Integration

Access gpt-oss-20b (High) via AnyAPI.ai

Access gpt-oss-20b (High) through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate gpt-oss-20b (High) without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access gpt-oss-20b (High) and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test gpt-oss-20b (High) against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use gpt-oss-20b (High) from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use gpt-oss-20b (High) for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

gpt-oss-20B supports a 131,072-token context window and generates up to 32,768 output tokens on managed APIs. The output ceiling is much smaller than the input capacity, so long single generations must be planned around the 32K limit rather than the full 128K context.

Yes. gpt-oss-20B is an open-weight model released by OpenAI under the permissive Apache 2.0 license, allowing commercial use, customization, and self-hosting without copyleft restrictions. OpenAI publishes the weights, tokenizer, and inference implementations, and the model is fully fine-tunable.

OpenAI states gpt-oss-20B delivers results similar to o3-mini on common benchmarks, matching or exceeding it on competition mathematics and health while being open-weight and far cheaper to run. However, o3-mini has a larger 200K context and higher output ceiling and leads on several proprietary benchmarks like GPQA and MMLU.

Yes. Thanks to its Mixture-of-Experts design and native MXFP4 quantization, gpt-oss-20B fits within roughly 16 GB of memory and runs on consumer GPUs or capable laptops. This makes it suitable for on-device inference, offline agents, and privacy-sensitive deployments without data-center hardware.

Yes. gpt-oss-20B has native function calling plus web browsing and Python execution tooling, and supports Structured Outputs. It exposes configurable reasoning effort—low, medium, or high—set in the system prompt, letting you trade latency for accuracy. It also provides full chain-of-thought visibility for debugging.

* Benchmark data source: Artificial Analysis artificialanalysis.ai