A delay blamed on the reward
Grok 4.7 is not shipping yet. Elon Musk said in a reply on X on 11 September that the model “needs a few more days to cook”, and gave an unusually technical reason: SpaceXAI, the company behind Grok, might have penalised response length too much during reinforcement learning.
The result, in Musk’s description, is a model that still “gives up on hard tasks (that it can do!) too early” and is not yet rigorous enough in checking its own work. He hedged the diagnosis himself, adding “or something” to the length explanation, and gave no new release date. TeslaNorth noted that Musk framed the hold-up as a tuning problem rather than a missing feature or a hardware limit; Data Studios also reported the delay.
A date that has moved before
Musk has given public timelines for the model since July. According to a compilation of his posts by CellCog, he wrote on 24 July that Grok 4.7 was four weeks away, on 12 August that it should be ready in three to four weeks, and at the start of September that it “comes out in 10 days”. That last post, published in the early hours of 2 September European time, pointed to a launch around 12 September, as Big Hat Group and NextBigFuture reported.
Musk has also described Grok 4.7 as a 2.1-trillion-parameter model and, according to NextBigFuture, said it would be better than Grok 4.6 in every way except serving speed, with better token efficiency. None of that comes from the company itself: SpaceXAI’s developer release notes contain no Grok 4.7 entry, model ID or price.
What SpaceXAI last shipped
The current model is Grok 4.6, released on 12 August. SpaceXAI said it builds on Grok 4.5, which arrived in July, with a particular focus on long-running agents. The release notes list a 500,000-token context window and pricing of $2 per million input tokens, $0.50 for cached input and $6 for output.
Long-running agents are also what the company is selling to businesses. On 3 September it opened Grok Bot to enterprises, adding access, network and audit controls and offering Grok and Cursor Enterprise customers two weeks of free organisation-wide use. A model that abandons hard work early is precisely the failure that pitch cannot afford.

What a length penalty is for
Reasoning models do better on hard problems by generating more tokens, and that verbosity wastes compute on easy ones. One established remedy is to add a length penalty to the reinforcement learning reward, so the model learns to answer more briefly.
Research shows how that can backfire. A May 2025 paper on stable reinforcement learning for efficient reasoning found that length-penalty rewards can make accuracy collapse as answers get shorter. Its method, GRPO-λ, drops the penalty when a model’s answers to a question are mostly wrong and keeps it when they are mostly right.
A June 2025 paper, Just Enough Thinking, argues against uniform penalties that treat every problem alike. Its adaptive penalty scales inversely with how often the model solves a prompt, so easy prompts pay a high cost for extra tokens while hard ones are left unhindered. On a 1.5-billion-parameter model, the authors reported cutting average token use by 50% without a significant drop in performance.

Why the explanation matters
Musk’s account fits that failure mode closely. A model charged for every token has a reason to stop before a hard problem is solved, and checking an answer costs tokens too. Neither Musk nor SpaceXAI has said what penalty was used or how it will change, so this remains an informal diagnosis rather than a technical disclosure.
It is still a specific one. It points at a reward-design choice that trades inference cost against persistence, the trade-off that research on efficient reasoning has been working through since at least 2025.
What to watch
The next concrete signal is a Grok 4.7 entry in SpaceXAI’s release notes, with a model ID, price and context window. If the length penalty was the problem, the evidence will be in how the model’s token use on hard tasks compares with Grok 4.6, and whether SpaceXAI’s own benchmark figures, when they appear, hold up in independent testing.