OpenAI's compact open-weight reasoning model that runs on 16 GB hardware with configurable reasoning and native tool calling.
Output Speed *
Intelligence Index *
Context Window *
Input price
Output price
Performance
Reasoning Density in a 16 GB Footprint
Benchmarks
Where gpt-oss-20B Lands in Independent Testing
Output Speed
Intelligence Index
MMLU *
GPQA *
HLE *
LiveCodeBench *
Technical Specifications
What the model supports
Quickstart
import requests
url = "https://api.anyapi.ai/v1/chat/completions"
payload = {
"model": "gpt-oss-20b",
"messages": [
{
"role": "user",
"content": "Hello"
}
]
}
headers = {
"Authorization": "Bearer AnyAPI_API_KEY",
"Content-Type": "application/json"
}
response = requests.post(url, json=payload, headers=headers)
print(response.json())const url = 'https://api.anyapi.ai/v1/chat/completions';
const options = {
method: 'POST',
headers: {Authorization: 'Bearer AnyAPI_API_KEY', 'Content-Type': 'application/json'},
body: '{"model":"gpt-oss-20b","messages":[{"role":"user","content":"Hello"}]}'
};
try {
const response = await fetch(url, options);
const data = await response.json();
console.log(data);
} catch (error) {
console.error(error);
}curl --request POST \
--url https://api.anyapi.ai/v1/chat/completions \
--header 'Authorization: Bearer AnyAPI_API_KEY' \
--header 'Content-Type: application/json' \
--data '{
"model": "gpt-oss-20b",
"messages": [
{
"role": "user",
"content": "Hello"
}
]
}'Limitations & Trade-offs
Best-Fit Workloads
Where this model earns its place
Local and edge reasoning agents
gpt-oss-20B's MoE sparsity and MXFP4 quantization let it run within roughly 16 GB of memory, enabling on-device agents on laptops or single consumer GPUs. Combined with native function calling and full chain-of-thought, it powers offline or privacy-sensitive assistants without cloud round-trips. Plan around the 32K output cap and provider-dependent latency.
Cost-sensitive high-volume tool-calling
With only 3.6B active parameters, gpt-oss-20B is inexpensive to serve at scale, making it a strong default for high-throughput agentic pipelines that issue many function calls. OpenAI positions the gpt-oss family for strong tool use, and it supports web browsing, Python execution, and developer-defined functions. Use low or medium reasoning effort to control per-call cost.
Fine-tuned domain reasoning
The Apache 2.0 license and full fine-tunability let teams adapt gpt-oss-20B to specialized domains without copyleft or patent risk. Researchers already LoRA-tune it for long-context and reasoning tasks. This suits companies needing a customizable, self-hostable reasoning model with predictable licensing, where a proprietary API would be restrictive or costly.
Structured extraction and workflow automation
Native Structured Outputs and JSON-schema support make gpt-oss-20B suitable for turning unstructured text into typed records within automation pipelines. Its o3-mini-class reasoning handles multi-field extraction and conditional logic, while configurable reasoning effort keeps routine extractions fast. Keep individual outputs under the 32K-token ceiling for large batch jobs.