Grok's new speech model cut short-phrase word error rate from 20.6% to 6.8%
SpaceXAI kept batch transcription at $0.10 an hour, and version 1.0 goes away in the coming weeks.
On September 18, 2026, SpaceXAI released Grok Voice Transcribe 2.0, and left the price where it was. Across its own real-world evaluations the company says the new model is twice as accurate as version 1.0.
The multilingual jump is the big one. On short voice-assistant phrases across 19 languages, word error rate falls from 20.6% to 6.8%.
- Released
- Sep 18, 2026
- From
- SpaceXAI
- Batch
- $0.10 per hour of audio
- Streaming
- $0.20 per hour
- Public ranking
- First for accuracy among 32 streaming models
Existing users don't have to do anything. SpaceXAI says current Speech-to-Text API integrations pick up the accuracy improvement with no code changes, and that surprised me for a whole version number.
What it was built for

Clean audio with one speaker is the easy case, and SpaceXAI says most models handle it.
Real-world audio is harder: flaky phone lines, competing voices, local accents, and phone numbers or email addresses read aloud.
It sits on the audio foundation model behind Grok Voice, which SpaceXAI says already runs tens of thousands of customer-support calls a day and transcribes millions of hours of video narration. The training audio is live, noisy and multilingual, recorded across a range of environments and then refined with post-training.
So the four internal sets are drawn from production traffic: customer-support calls over 8 kHz telephony, conversations with Grok, spoken credentials like account codes and email addresses, and short commands in 19 languages. Version 2.0 beats 1.0 on all four, and on telephony SpaceXAI says it leads every model it tested. Granted, those four sets are its own. The one outside number is the public Artificial Analysis leaderboard, where it ranks first for accuracy among 32 streaming models.
It handles dozens of languages, picks the language itself, and follows a switch mid-recording in a single pass. Short commands are the hard case there. A few words in a car give the model almost nothing to work out which language it's hearing.
What comes with it
Batch and streaming both, word-level timestamps with confidence scores, speaker labels at no extra cost, up to 8 channels transcribed separately, and up to 100 domain terms per request to bias it toward your product names or medical vocabulary. It also strips the ums and uhs if you want, and detects when a speaker has finished a turn.
Multilingual accuracy is its largest improvement over Grok Voice Transcribe 1.0
Atlassian has it behind Loom, the screen recorder, after finding it more accurate than what Loom was using (Atlassian's own comparison, on its own videos). Sanchan Saxena, an Atlassian SVP, describes the workflow it opens up as recording an action plan and piping the transcript into Cursor to make the code changes.
With Grok powering Loom's speech-to-text and Cursor turning that into code, we're closing the loop from context to code: record what you mean, and the work gets done.
Two rivals turn up on the word error rate chart, ElevenLabs Scribe v2 and Deepgram Nova-3, and the page draws those bars without printing what either one scored. (The accuracy-versus-price chart is the same.) The 20.6% and the 6.8% are the only word error rates written out anywhere in the post.
Grok Voice Transcribe 1.0 will be deprecated in the coming weeks. Pin grok-voice-transcribe-1.0 to stay on it through the transition.