The day after agreeing AI should slow down, Musk said Grok 4.8 finishes training this week
Grok 4.7 still isn't out, Grok 5 is his answer for AGI, and neither OpenAI nor Anthropic lists a size to set 2.5T against.
On September 12, 2026, Dario Amodei asked the AI industry to slow down, and Elon Musk replied an hour later: "Dario is right." The next night, answering a question about what xAI would ship in September, he posted that Grok 4.8 is a 2.5T model trained on a new C++ software stack. It "will finish training this week and start RL", he wrote.
I kept coming back to that turn. Musk agreed in public that the frontier should be paced, then spent the next two days telling people which Grok comes next.
- Musk agrees with Amodei
- Musk puts Grok 4.8 at 2.5T
- Musk names Grok 5 for AGI
- Musk ranks Grok 4.7 to 5
Eastern time, read off each post on X.
- Grok 4.8
- A 2.5T model
- Stage
- Finish training this week and start RL
- Grok 4.7
- Needs a few more days to cook
- AGI
- That will be Grok 5
What 2.5T means
The T is trillions of parameters, which is the size of the model. It isn't tokens (the unit context windows are counted in), so it says nothing about how much text Grok 4.8 can read at once.

That makes it kind of hard to set against the models Musk compares Grok with. GPT-6 Astra and Claude Fable 5.1 both come with published context windows. Neither lab's model page gives a parameter count.
- GPT-6 Astra context
- 1,050,000 tokens
- Claude Fable 5.1 context
- 1M tokens
- DeepSeek-V4-Pro size
- 1.7T params
- Largest Qwen3.8
- 2.4T, A95B
The open-weight labs do publish sizes. DeepSeek's Hugging Face page lists DeepSeek-V4-Pro at 1.7T parameters, and Alibaba puts the size of its largest Qwen3.8 right in the name, Qwen3.8-2.4T-A95B. By Musk's figure, Grok 4.8 is bigger than both.
Where Musk ranks it
The morning after the 2.5T post, a user predicted Grok 4.7 would land in the same category as Fable 5.1 and Astra. Musk answered with the whole order.
Agree with “PACE THE FRONTIER” in public. Keep training. The 2.5T Grok 4.8 model is going to be better than the current non-public models from Anthropic & OpenAI. I think Grok 4.7 will be in the same category as Fable 5.1 & Astra. That’s probably why he’s delayed the release twice. And the SpaceXAI engineers are in full monk mode. OpenAI is 11 years old. Anthropic is 5. SpaceXAI is 3. Elon has done a hell of a job closing the gap with labs that had years more experience. And this guy isn’t slowing down.
Grok 4.7 should be roughly on par with Opus 5.0, not 5.1. Better in some ways, worse in others. We need to fix multimodal performance. Grok 4.8 will be a noticeable improvement. Grok 4.9 is probably Astra/Fable class. Grok 5 maybe better than anything. We shall see.
By his own ordering, the one he expects to reach the Astra and Fable class is Grok 4.9. Grok 4.8 sits below it.
One line needs a footnote. There's no Opus 5.1 on Anthropic's model list, which runs Claude Fable 5.1, Claude Opus 5, Claude Sonnet 5 and Claude Haiku 4.5. My read (and I could be wrong) is that "Opus 5.0, not 5.1" means Opus 5 and not Fable 5.1. Our Opus 5.1 page explains the name.
What xAI hasn't said
None of this is on xAI's own site. Its news page has no Grok 4.7 or Grok 4.8 post, so there's no release date from the lab itself. No price or benchmark either. Grok 4.8 and Grok 4.7 have pages here now, marked announced, and they'll fill in when xAI publishes.
Grok 4.7 needs a few more days to cook. We might have penalized response length too much (or something) in RL, as it still gives up on hard tasks (that it can do!) too early and isn’t yet sufficiently rigorous in checking its work.
That was September 11, before any of the Grok 4.8 posts. A few more days is the first estimate worth holding him to.