xAI's own table puts Grok 4.7 ahead of Fable 5.1 on three of seven benchmarks
Input tokens cost $2 a million against Fable 5.1 Max's $10, and output $6 against $50, on a table xAI drew itself.
SpaceXAI put out Grok 4.7 on September 21, 2026, and it was in Cursor and Grok Build the same day. xAI calls it its most capable model for coding and knowledge work.
It runs at Grok 4.6's price and speed. Input tokens are $2 per million and output $6, and on the same table Fable 5.1 Max is $10 and $50 (those are xAI's figures for a competitor, so take them as xAI's).
- Model
- Grok 4.7, from SpaceXAI
- Input price
- $2 per million input tokens
- Output price
- $6 per million output tokens
- Where
- Cursor, Grok Build, the Grok API, model routers and cloud platforms
What changed since 4.6
Grok 4.7 uses a new, larger base model. xAI says it trained the model on a harder mix of tasks and pushed it toward the long ones, and that the result is better at checking its own work and at managing longer context. It was also trained to natively understand the Grok Bot harness, which is xAI's way of saying the chat product and the model now know each other.
It was trained with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete.
The scores xAI printed
Seven benchmarks are on the table with Fable 5.1 Max and GPT-5.6 Sol Max beside them. Grok 4.7 comes out ahead of Fable 5.1 on electrical engineering, 64.0% to 56.4%, and on Harvey's legal agent benchmark, 19.6% to 6.7%. It edges DeepSWE v1.1 at 71.0% against 70.0%, with an asterisk saying that score is high effort.
Fable 5.1 keeps the rest. CursorBench 4.0 goes 51.8% to 46.3%, clinical reasoning 62.1% to 56.7%, and Terminal-Bench 4.0 is not close at all: 57.9% against 38.0%. So the cheap model wins three of seven and loses the one about working in a terminal for hours, which is the job people buy a coding model to do.
Twice as fast, at half the price of comparable models
Harvey's legal agent benchmark is the row where nobody does well. GPT-5.6 Sol Max gets 2.5% on it, and Grok 4.7's 19.6% is the best score on the table. xAI printed the whole row anyway, and that surprised me more than the win did.
Safeguards, and who gets the red-team access
xAI says Grok 4.7 was built with an entirely new safeguard stack, and that it's the strongest model the company has tested on refusals and jailbreak resistance. It tops LatchBio's biosafety benchmark at 62.4% (xAI doesn't say who else is on that leaderboard). On HackerBench v0.3, its own benchmark for risky and malicious cyber tasks, the model lets 3.3% of risky dual-use prompts through.
There's a faster variant too, at twice the output speed for twice the price. xAI has also started handing select cybersecurity partners invite-only access to Grok 4.7's red-team capabilities for defense research, and it doesn't name any of them.