ThinkFacility

News

The day after agreeing AI should slow down, Musk said Grok 4.8 finishes training this week

Grok 4.7 still isn't out, Grok 5 is his answer for AGI, and neither OpenAI nor Anthropic lists a size to set 2.5T against.

On September 12, 2026, Dario Amodei asked the AI industry to slow down, and Elon Musk replied an hour later: "Dario is right." The next night, answering a question about what xAI would ship in September, he posted that Grok 4.8 is a 2.5T model trained on a new C++ software stack. It "will finish training this week and start RL", he wrote.

I kept coming back to that turn. Musk agreed in public that the frontier should be paced, then spent the next two days telling people which Grok comes next.

  1. Musk agrees with Amodei
  2. Musk puts Grok 4.8 at 2.5T
  3. Musk names Grok 5 for AGI
  4. Musk ranks Grok 4.7 to 5

Eastern time, read off each post on X.

Grok 4.8
A 2.5T model
Stage
Finish training this week and start RL
Grok 4.7
Needs a few more days to cook
AGI
That will be Grok 5

What 2.5T means

The T is trillions of parameters, which is the size of the model. It isn't tokens (the unit context windows are counted in), so it says nothing about how much text Grok 4.8 can read at once.

Elon Musk seated at a conference table with microphones and water bottles, between a woman in a red jacket and a man in a blue suit
Elon Musk at a session on the risks of losing control of frontier AI, at the UK's AI Safety Summit at Bletchley Park in November 2023. Photo: UK Government, CC BY 2.0, via Wikimedia Commons

That makes it kind of hard to set against the models Musk compares Grok with. GPT-6 Astra and Claude Fable 5.1 both come with published context windows. Neither lab's model page gives a parameter count.

GPT-6 Astra context
1,050,000 tokens
Claude Fable 5.1 context
1M tokens
DeepSeek-V4-Pro size
1.7T params
Largest Qwen3.8
2.4T, A95B

The open-weight labs do publish sizes. DeepSeek's Hugging Face page lists DeepSeek-V4-Pro at 1.7T parameters, and Alibaba puts the size of its largest Qwen3.8 right in the name, Qwen3.8-2.4T-A95B. By Musk's figure, Grok 4.8 is bigger than both.

Where Musk ranks it

The morning after the 2.5T post, a user predicted Grok 4.7 would land in the same category as Fable 5.1 and Astra. Musk answered with the whole order.

Lueul@itslueul

Agree with “PACE THE FRONTIER” in public. Keep training. The 2.5T Grok 4.8 model is going to be better than the current non-public models from Anthropic & OpenAI. I think Grok 4.7 will be in the same category as Fable 5.1 & Astra. That’s probably why he’s delayed the release twice. And the SpaceXAI engineers are in full monk mode. OpenAI is 11 years old. Anthropic is 5. SpaceXAI is 3. Elon has done a hell of a job closing the gap with labs that had years more experience. And this guy isn’t slowing down.

View the post on X
Elon Musk@elonmusk · xAI

Grok 4.7 should be roughly on par with Opus 5.0, not 5.1. Better in some ways, worse in others. We need to fix multimodal performance. Grok 4.8 will be a noticeable improvement. Grok 4.9 is probably Astra/Fable class. Grok 5 maybe better than anything. We shall see.

View the post on X

By his own ordering, the one he expects to reach the Astra and Fable class is Grok 4.9. Grok 4.8 sits below it.

One line needs a footnote. There's no Opus 5.1 on Anthropic's model list, which runs Claude Fable 5.1, Claude Opus 5, Claude Sonnet 5 and Claude Haiku 4.5. My read (and I could be wrong) is that "Opus 5.0, not 5.1" means Opus 5 and not Fable 5.1. Our Opus 5.1 page explains the name.

What xAI hasn't said

None of this is on xAI's own site. Its news page has no Grok 4.7 or Grok 4.8 post, so there's no release date from the lab itself. No price or benchmark either. Grok 4.8 and Grok 4.7 have pages here now, marked announced, and they'll fill in when xAI publishes.

Elon Musk@elonmusk · xAI

Grok 4.7 needs a few more days to cook. We might have penalized response length too much (or something) in RL, as it still gives up on hard tasks (that it can do!) too early and isn’t yet sufficiently rigorous in checking its work.

View the post on X

That was September 11, before any of the Grok 4.8 posts. A few more days is the first estimate worth holding him to.