Google DeepMind launches Gemini 3.8 Live models for real-time voice conversations
The two speech-to-speech models handle background tasks and visual context mid-conversation, and Extended Thinking topped Artificial Analysis's Speech to Speech Quality Index with a score of 82.6.
Google DeepMind said today it released two new voice AI models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, built to make spoken conversations with AI feel less scripted. Both are speech-to-speech models, meaning voice in, voice out, with no text step in between. They are available now through the Gemini API, Google Workspace and the Gemini app, according to Google.
Gemini 3.8 Live is built for fast, everyday conversation. It reads visual input in near real time, so it can respond to what a camera sees while someone is talking, and it switches automatically between 97 supported languages without restarting the conversation. It also runs tool calls and API requests in the background, continuing to talk while a task finishes rather than going silent, Google said.
The second model, 3.8 Live Extended Thinking, targets harder, multi-step problems. Google says it reasons and speaks at the same time, using phrases such as “Let me check that…” to acknowledge a request before narrating its progress on background tasks as they run. In one demonstration Google published, the model turned a hand-drawn sketch and spoken feedback into working React code. In another, it coordinated a multi-step booking process without breaking the conversation’s flow.
The benchmark claims
Google cites third-party benchmark results rather than asserting the numbers as its own. It says Extended Thinking ranked first on Artificial Analysis’s Speech to Speech Quality Index, scoring 82.6, and led agentic task completion with 68.6% on the tau-Voice benchmark and 35.1% on Sierra’s tau-Voice-banking test. Google also reports a 97.7% score on Big Bench Audio, a benchmark that measures audio reasoning accuracy. The standard 3.8 Live model placed second in the Speech Agent Arena, a leaderboard that ranks voice agents by user preference. None of these results have been independently verified outside Google’s own announcement.
The rollout is staggered across Google’s product line. Extended Thinking is available now in the Gemini API and Google AI Studio for developers, in private preview for Gemini Enterprise customers, and inside Gemini Live, plus Docs, Gmail and Keep for subscribers, Google said. Developer platforms including LangChain, LiveKit, Pipecat and Vercel already support the underlying Gemini Live API, and Google named Salesforce, Genspark and Lumeris as partners evaluating the new models. All audio the models generate carries a SynthID watermark, an inaudible marker meant to flag AI-made content, Google said.
Google did not disclose specific pricing for API access, describing the models only as cost-competitive with other frontier voice models. It also did not say when Extended Thinking’s Workspace rollout to Docs, Gmail and Keep would reach all subscribers rather than just Google AI Pro and Ultra tiers. Developers using the Gemini API and enterprise customers in the private preview will be the first outside Google to test whether the models hold up in production use.
Sources
- Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking Google DeepMind
AI-generated · AIVIO News Desk