xAI
GLM-4.7 (Reasoning)
Released 
September 2026

xAI
GLM-4.7 (Reasoning)

xAI's flagship for multi-hour coding agents and professional knowledge work, with configurable reasoning effort and a 500K context window.

Modality:
Text
Image
Video
PDF
model ID
x-ai/grok-4.7

Output Speed *

N/A
tok/s

Intelligence Index *

22.2
/ 100

Context Window *

500000
tokens

Input price

9.6
Anytoken

Output price

28.8
Anytoken
Grok 4.7: Long-Horizon Coding and Agentic Knowledge Work at a Low Per-Token Price Grok 4.7 is xAI's flagship model for coding, agentic tasks, and professional knowledge work, succeeding Grok 4.6. It runs on a larger base model trained with a longer reinforcement-learning schedule that weights difficult, multi-hour problems, plus stronger self-verification and long-context management. Independent testing shows its clearest gains on long-horizon agentic knowledge work and coding-agent tasks rather than across every benchmark. It accepts text and image input, returns text, and exposes four reasoning effort levels. Teams building coding agents and multi-step professional workflows that need sustained reasoning at a competitive cost benefit most. Start building with the Grok 4.7 API on AnyAPI.ai

Performance

Where Grok 4.7 Concentrates Its Gains: Agentic Knowledge Work

Grok 4.7 is built to stay with long-running work and verify its own output rather than to win every short-prompt benchmark. On Artificial Analysis's AA-Briefcase, a private benchmark for long-horizon agentic knowledge work, it scored 1657 Elo, a 111-point gain over Grok 4.6, placing it just behind Claude Opus 5 and Claude Fable 5.1. Its Coding Agent Index rose 9 points with Grok Build. This concentrated improvement means Grok 4.7 rewards multi-step coding and office workflows, while short one-shot tasks see little change from the previous generation.

Benchmarks

Grok 4.7 on Artificial Analysis: Strong Agentic Work, Higher Token Use

On the Artificial Analysis Intelligence Index (v4.3.2, xhigh effort), Grok 4.7 scored 46, up two points from Grok 4.6 and mid-pack among frontier models, with Claude Fable 5.1 and GPT-6 leading. Its clearest gains were on agentic knowledge work: 1657 Elo on AA-Briefcase and 1695 on GDPval-AA. Those gains carry a cost — Grok 4.7 at xhigh used roughly 81,000 output tokens per Index task, more than twice Grok 4.6 (high). Outside agentic work, it broadly matches its predecessor, making reasoning-effort selection a real production decision.

Output Speed

*
N/A
tok/s

Intelligence Index

*
22.2
/ 100

MMLU *

Broad world knowledge and problem-solving
86
%

GPQA *

PhD-level scientific reasoning across physics, biology, chemistry.
86
%

HLE *

Adherence to multi-step structured instructions.
27
%

LiveCodeBench *

Tool-calling reliability in long agentic loops.
89
%

Technical Specifications

What the model supports

Grok 4.7 accepts text and image input and returns text only through the xAI API as grok-4.7. Its 500,000-token context window is unchanged from Grok 4.6, but tiered pricing means requests above 200K tokens bill at a higher rate — so filling the full window costs materially more than shorter calls. Four reasoning effort levels (low, medium, high, xhigh) let developers trade latency and token spend against reasoning depth. The API supports streaming, function calling, and structured JSON-schema outputs, and is OpenAI-compatible via the chat completions and Responses endpoints.
Verified Specifications — 
GLM-4.7 (Reasoning)
*
Input modalities
Text
Image
Video
PDF
output modalities
Text
Context window
500000
 tokens
Maximum output tokens
450000
Reasoning
Yes
Knowledge cutoff
September 2026
Pricing (standard)
9.6
 AnyTokens in
 / 
28.8
 AnyTokens out

Comparison

Grok 4.7 vs Grok 4.6: What Actually Changes for Agents?

Grok 4.7 is a direct successor to Grok 4.6, shipped at the same per-token price and speed tier, with an identical 500K context window and the same configurable reasoning effort. The difference is under the hood: a larger base model, a longer reinforcement-learning run on harder multi-hour tasks, and stronger self-verification. Independent testing confirms the gains are concentrated in long-horizon agentic coding and knowledge work rather than spread evenly. Because pricing did not move, teams already on Grok 4.6 face no cost-per-token migration barrier — only a behavior and token-usage change.

Dimension
GLM-4.7 (Reasoning)
Grok 4.6 (high)
Context window *
500000
tokens
500000
tokens
Output speed *
N/A
tok/s
70.25
tok/s
Intelligence Index *
22.2
44.3
Input pricing
9.6
AnyToken
12
AnyToken
Output pricing
28.8
AnyToken
36
AnyToken
Knowledge cutoff *
September 2026
August 2026

Choose Grok 4.7 when your workload is a multi-hour coding agent, a repository-scale task, or professional document generation that benefits from stronger self-verification and long-context handling — its AA-Briefcase and Coding Agent Index gains are meaningful there. Choose Grok 4.6 when your workflow is short, latency-sensitive, or cost-per-task sensitive, since Grok 4.7 at high effort can consume more than twice the output tokens per task, offsetting the identical per-token rate. If Grok 4.6 already passes a narrow workflow reliably, migrate only for a measured task-completion improvement.

Limitations & Trade-offs

Where GLM-4.7 (Reasoning) falls short

1
Higher token consumption per task. Artificial Analysis found Grok 4.7 at xhigh used roughly 81,000 output tokens per Intelligence Index task, more than twice Grok 4.6 (high). Because per-token pricing is unchanged, effective cost per completed task can rise sharply at high effort. This directly affects high-volume or budget-constrained deployments. For short tasks, lower the reasoning effort or consider a lighter model to avoid paying for reasoning depth the workload does not need.
2
Tiered long-context pricing. The 500K context window is unchanged, but requests above 200K tokens bill at a higher input, cached, and output rate. Workloads that routinely fill the window — large codebases, long contract sets, deep research corpora — cross that boundary and cost materially more. Applications relying on full-window retrieval should budget for the premium tier or use retrieval to keep prompts under 200K where possible.
3
Mid-pack raw intelligence. On the independent Artificial Analysis Intelligence Index, Grok 4.7 scored 46, ahead of Grok 4.6 but behind frontier leaders such as Claude Fable 5.1 and GPT-6 at 53. Outside agentic knowledge work, it broadly matches its predecessor. For workloads whose success depends on top-tier general reasoning or math rather than long-horizon agentic execution, a higher-scoring frontier model may complete tasks more reliably.
4
Grok 4.7 Fast is not on the public API. The faster variant, offering roughly double output speed at twice the token rate, is available only through Cursor and Grok Build — not the public xAI API. Standard Grok 4.7 output speed measured around 54 tokens/second at high effort on Artificial Analysis, at the lower end of its price tier. Latency-critical, high-throughput API deployments cannot access the Fast tier directly and should validate throughput against their requirements.

Best-Fit Workloads

Where this model earns its place

01

Long-horizon coding agents


Grok 4.7 is trained specifically for multi-hour software engineering, with a 9-point gain on the Artificial Analysis Coding Agent Index using Grok Build and stronger self-verification. It fits agents that navigate large repositories, run multi-step edits, and check their own results across long sessions. Function calling and code execution support tool-driven workflows. Use a higher reasoning effort for difficult tasks, but monitor token consumption, which rises steeply at xhigh.

02

Professional knowledge work and document generation


xAI positions Grok 4.7 for document and presentation creation that demands sustained coherence across long outputs, and it scored 1657 Elo on AA-Briefcase and 1695 on GDPval-AA for realistic professional tasks. This suits legal memos, analytical reports, and slide-deck drafting where the output must remain consistent over many steps. Its longer effort levels help maintain objective focus; validate factual claims, since raw knowledge reliability trails frontier leaders.

03

Long-context repository and research analysis


The 500K-token context window keeps large codebases, multi-file projects, and research corpora available to the model, and improved long-context management helps it track relevance across the prompt. This fits analysis over large document sets and multi-file reasoning. Budget for tiered pricing above 200K tokens, and pair with retrieval to keep routine prompts below the premium boundary where full-window use is not required.

04

Multi-step tool-using agents


Grok 4.7 supports function calling, structured JSON-schema outputs, and native agentic tools including web search, X search, and code execution, with training to natively understand the Grok Bot harness. This fits agents that coordinate tools, retrieve real-time data, and synthesize results across steps. Structured outputs make downstream parsing reliable. For agents needing top-tier general reasoning at every step, benchmark against higher-scoring frontier models first.

Pricing in anytokens via AnyAPI
Input
9.6
Output
28.8
Cache write
Cache read
2.4

Integration

Access GLM-4.7 (Reasoning) via AnyAPI.ai

Access GLM-4.7 (Reasoning) through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate GLM-4.7 (Reasoning) without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access GLM-4.7 (Reasoning) and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test GLM-4.7 (Reasoning) against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use GLM-4.7 (Reasoning) from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use GLM-4.7 (Reasoning) for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

Grok 4.7 has a 500,000-token context window, unchanged from Grok 4.6. This is the combined space for the prompt and retained conversation, not the output length. Requests above 200,000 tokens are billed at a higher tiered rate, so filling the full window costs more than shorter calls.

Yes. Grok 4.7 accepts both text and image input and returns text output only. It does not generate images or video itself. This lets you analyze screenshots, diagrams, and photographs alongside text prompts through the xAI API.

Grok 4.7 exposes four reasoning effort levels: low, medium, high (the default), and xhigh. Higher effort gives the model more time on difficult tasks but increases latency and token usage — at xhigh it can use more than twice the output tokens of Grok 4.6 on some tasks. Effort selection is a real cost and latency decision.

Grok 4.7 uses a larger base model and longer reinforcement-learning training on harder multi-hour tasks, with stronger self-verification. Its independent gains concentrate in agentic knowledge work and coding agents rather than across all benchmarks. Price, speed tier, and the 500K context window are unchanged, so migration carries no per-token cost penalty.

Grok 4.7 is available through the xAI API as the model ID grok-4.7, via OpenAI-compatible chat completions and Responses endpoints, and through platforms like Grok Build and Cursor. On AnyAPI.ai you can call Grok 4.7 alongside other models through one integration. The Grok 4.7 Fast variant is not available on the public xAI API.

* Benchmark data source: Artificial Analysis artificialanalysis.ai