A bigger base model at the old price
SpaceXAI released Grok 4.7 on Monday, and kept the price of its predecessor. Input costs $2 per million tokens and output $6 per million, the same as Grok 4.6, with a fast variant served at twice the output speed for twice the price. The company describes it as its most capable model for coding and knowledge work.
The changes underneath are structural rather than cosmetic. SpaceXAI says Grok 4.7 uses a new, larger base model than Grok 4.6, trained with a longer reinforcement learning run on a harder mix of tasks weighted towards problems that take many hours. It also says the model was trained to natively understand the Grok Bot harness, its own agent runtime.
The company’s numbers, and someone else’s
SpaceXAI’s own table puts Grok 4.7 at 46.3% on CursorBench 4.0, up from 40.4% for Grok 4.6, and ahead of the 41.7% it records for GPT-5.6 Sol but behind Claude Fable 5.1 at 51.8%. On EEBench, an electrical engineering test, it reports 64.0% against 39.4% for GPT-5.6 Sol and 56.4% for Fable 5.1. On the Harvey legal agent benchmark it reports 19.6%, against 6.7% for Fable 5.1.

Those are vendor-reported figures. The independent picture is less flattering. Artificial Analysis, which runs its own evaluations, scored Grok 4.7 at 46 on its Intelligence Index v4.3.2, against 53 for both Claude Fable 5.1 and GPT-6, according to The Decoder.
The gap is clearest on one test both sides ran. SpaceXAI reports 38.0% on Terminal-Bench 4.0. Artificial Analysis measured 26% — roughly where it puts DeepSeek V4.1 Flash, at 27%, and well below the 60% it records for GPT-6 Astra and 55% for Claude Fable 5.1.
Where it does lead
On GDPval, which scores professional work, SpaceXAI publishes an Elo of 1,695 for Grok 4.7 at its highest effort setting, behind Fable 5.1 at 1,735 but ahead of GPT-6 Astra at 1,542. The company also reports 1,657 on AA Briefcase v1.1, just under Fable 5.1’s 1,678.

Read together, the vendor table and the independent index tell a consistent story about positioning rather than a contradiction: the model does well on multi-hour, document-shaped professional work and less well on long-running terminal tasks, at a price closer to Chinese open-weight models than to Western frontier ones.
A new safeguard stack
SpaceXAI says Grok 4.7 ships with an entirely new safeguard stack and is the strongest model it has tested on refusals and jailbreak resistance. It reports topping LatchBio’s biosafety benchmark at 62.4%, and says the model allows 3.3% of risky dual-use prompts through on HackerBench v0.3, its own internal cyber test. Both of those are the company’s own measurements.
It has also begun giving selected cybersecurity partners invite-only access to Grok 4.7’s red-team capabilities for defence research — a narrower version of the access question that has divided the labs all year.
What to watch
The model is available now in Cursor and Grok Build, through the Grok API, and via third-party coding harnesses, model routers and cloud platforms. The thing to watch is whether the Terminal-Bench discrepancy narrows as more independent evaluations land, and whether $2 per million input tokens is enough to move developers who are currently paying $10 for Fable 5.1.