xAI releases Grok 4.7, its new frontier model for coding and agentic work
xAI shipped Grok 4.7 on September 21, 2026, calling it "our most capable model for coding and knowledge work" and pitching it as roughly twice as fast at half the price of comparable frontier models, while holding the same price and speed as its predecessor, Grok 4.6.
What's new
Grok 4.7 is built on a larger base model than Grok 4.6 and was trained with a longer reinforcement-learning run on a harder mix of tasks, weighted toward problems that take many hours to finish. xAI says the model is better at verifying its own work and managing longer context, and it was trained to natively understand the Grok Bot harness, improving conversational and general-knowledge performance alongside coding.
On xAI's own benchmark chart, Grok 4.7 scores 46.3% on CursorBench 4.0 (versus 40.4% for Grok 4.6), 71.0% on DeepSWE v1.1 at high effort, and 64.0% on EEBench, an electrical-engineering benchmark. It also posts gains on GDPval and AA Briefcase, benchmarks meant to approximate multi-hour professional office work such as legal, nursing, and financial-analyst tasks.
Per xAI's developer docs, Grok 4.7 is available on the API as grok-4.7 with a 500k-token context window, text-and-image input with text-only output, and four reasoning-effort levels (low, medium, high by default, and xhigh). Pricing is $2 per million input tokens, $0.50 for cached input, and $6 per million output tokens for prompts under 200k tokens, rising to $4/$1/$12 above that threshold — unchanged from Grok 4.6's rates.
On safety, xAI says Grok 4.7 was built with an entirely new safeguard stack and is "the strongest model we've tested on refusals and jailbreak resistance." The company highlights a 62.4% score on LatchBio's biosafety benchmark and says the model blocks only 3.3% of risky prompts on HackerBench v0.3, its benchmark for dual-use cybersecurity tasks, while rarely refusing legitimate security work. xAI is also giving select cybersecurity partners invite-only access to the model's red-team capabilities for defense research.
Context
Grok 4.7 follows Grok 4.6 and continues xAI's pattern of frequent, incremental frontier releases rather than large generational jumps. The announcement leans heavily on coding- and agentic-task benchmarks like CursorBench and Terminal-Bench, an emphasis that tracks the broader industry shift toward evaluating models on long-running, tool-using work rather than single-turn Q&A. xAI's own comparison chart places Grok 4.7 alongside Fable 5.1, Opus 5, GPT-5.6 Sol, Sonnet 5, and GPT-6 Astra, underscoring how crowded the frontier tier has become by late 2026.
The model launches directly inside Cursor and xAI's own Grok Build tool, in addition to the API and "third-party coding harnesses and cloud platforms," reflecting how much of the current competition among frontier labs plays out through coding-agent integrations rather than standalone chat products.
Why it matters
Holding pricing flat while claiming a meaningful CursorBench jump is xAI's most direct price-performance pitch yet to developers already spending heavily on coding agents. The heavier safety framing — a new safeguard stack, biosafety and cyber-dual-use benchmarks, and invite-only red-team access for security partners — also signals that xAI is trying to close a perception gap with rivals on model safety ahead of wider enterprise and government adoption, a theme it has leaned into as competitors publish their own safety and governance frameworks.