Google's new Gemini voice models talk you through a task while the tool calls run.
Google's own numbers put the better of the two at 35.1% on the hardest voice-agent benchmark it cites.
Google started rolling out two new voice models on September 15, 2026. Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking went to developers in the Gemini API and Google AI Studio, to everyone in Search Live, and to paying subscribers in the Gemini app and parts of Workspace.
Both are live dialogue models, and Google calls them "our most advanced live dialogue models yet". The narration is the part I didn't expect to see shipping.
- Models
- Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking
- Rolling out
- starting today in the Gemini API and Google AI Studio
- Languages
- 97 supported languages, mid-conversation
- Speech to speech
- 82.6, the #1 overall spot
- For everyone
- In Search Live
It keeps talking while it works

Extended Thinking "reasons and speaks simultaneously", the post says. It answers with a verbal cue while it's still thinking, then talks you through a multi-step job as the steps finish, so a long task doesn't leave the line silent.
It executes tools and API calls in the background while continuing the conversation, so the model can acknowledge requests and keep chatting while tasks finish in the background.
The cheaper of the two does the seeing. Google says 3.8 Live processes visual inputs in near real-time, and that it "automatically detects and transitions between 97 supported languages" mid-conversation. Two of the demos Google published are an employee onboarding walkthrough and a game of chess.
The plumbing for developers runs through the Gemini Live API, and Google names LangChain, LiveKit, Pipecat and Vercel among the platforms that handle the real-time media streaming underneath. Those are the pieces that keep an audio stream open while a model goes away and does something.
What the benchmarks say
Google puts Extended Thinking first on Artificial Analysis' Speech to Speech Quality Index at 82.6, and 3.8 Live second in the Speech Agent Arena. On Big Bench Audio, Extended Thinking scores 97.7%.
The agent scores sit a long way below that. Google says the model leads on agentic task completion with 68.6% on τ-Voice and 35.1% on Sierra's τ-Voice-banking benchmark. Finishing about a third of the banking tasks is the leading result, by Google's own account, which says something about where voice agents actually are.
Where you can use it
Search Live carries 3.8 Live to anyone. Extended Thinking goes into Gemini Live, into Docs for Google AI Pro and Ultra subscribers, and into Gmail and Keep for all Google AI subscribers. Enterprises get a private preview in Gemini Enterprise, with Gemini Enterprise for Customer Experience listed as coming soon.
Google didn't print a price, only that the models hold "a highly competitive price point compared to other frontier models". Every clip either model speaks carries a SynthID watermark, which Google says is "woven directly into the audio output".