xAI's flagship for multi-hour coding agents and professional knowledge work, with configurable reasoning effort and a 500K context window.
Output Speed *
Intelligence Index *
Context Window *
Input price
Output price
Performance
Where Grok 4.7 Concentrates Its Gains: Agentic Knowledge Work
Benchmarks
Grok 4.7 on Artificial Analysis: Strong Agentic Work, Higher Token Use
Output Speed
Intelligence Index
MMLU *
GPQA *
HLE *
LiveCodeBench *
Technical Specifications
What the model supports
Comparison
Grok 4.7 vs Grok 4.6: What Actually Changes for Agents?
Grok 4.7 is a direct successor to Grok 4.6, shipped at the same per-token price and speed tier, with an identical 500K context window and the same configurable reasoning effort. The difference is under the hood: a larger base model, a longer reinforcement-learning run on harder multi-hour tasks, and stronger self-verification. Independent testing confirms the gains are concentrated in long-horizon agentic coding and knowledge work rather than spread evenly. Because pricing did not move, teams already on Grok 4.6 face no cost-per-token migration barrier — only a behavior and token-usage change.
Choose Grok 4.7 when your workload is a multi-hour coding agent, a repository-scale task, or professional document generation that benefits from stronger self-verification and long-context handling — its AA-Briefcase and Coding Agent Index gains are meaningful there. Choose Grok 4.6 when your workflow is short, latency-sensitive, or cost-per-task sensitive, since Grok 4.7 at high effort can consume more than twice the output tokens per task, offsetting the identical per-token rate. If Grok 4.6 already passes a narrow workflow reliably, migrate only for a measured task-completion improvement.
Limitations & Trade-offs
Best-Fit Workloads
Where this model earns its place
Long-horizon coding agents
Grok 4.7 is trained specifically for multi-hour software engineering, with a 9-point gain on the Artificial Analysis Coding Agent Index using Grok Build and stronger self-verification. It fits agents that navigate large repositories, run multi-step edits, and check their own results across long sessions. Function calling and code execution support tool-driven workflows. Use a higher reasoning effort for difficult tasks, but monitor token consumption, which rises steeply at xhigh.
Professional knowledge work and document generation
xAI positions Grok 4.7 for document and presentation creation that demands sustained coherence across long outputs, and it scored 1657 Elo on AA-Briefcase and 1695 on GDPval-AA for realistic professional tasks. This suits legal memos, analytical reports, and slide-deck drafting where the output must remain consistent over many steps. Its longer effort levels help maintain objective focus; validate factual claims, since raw knowledge reliability trails frontier leaders.
Long-context repository and research analysis
The 500K-token context window keeps large codebases, multi-file projects, and research corpora available to the model, and improved long-context management helps it track relevance across the prompt. This fits analysis over large document sets and multi-file reasoning. Budget for tiered pricing above 200K tokens, and pair with retrieval to keep routine prompts below the premium boundary where full-window use is not required.
Multi-step tool-using agents
Grok 4.7 supports function calling, structured JSON-schema outputs, and native agentic tools including web search, X search, and code execution, with training to natively understand the Grok Bot harness. This fits agents that coordinate tools, retrieve real-time data, and synthesize results across steps. Structured outputs make downstream parsing reliable. For agents needing top-tier general reasoning at every step, benchmark against higher-scoring frontier models first.