Two models, one Live API
Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on Tuesday, two native speech-to-speech models built for production voice agents rather than for chat. Both are available now through the Gemini Live API and Google AI Studio. Gemini Enterprise gets them in private preview, and Google says they are rolling into Search Live and the Gemini Live app.
The pair splits along cost. Gemini 3.8 Live is the cheaper conversational model, pitched at fluid dialogue and visual grounding. Extended Thinking is the variant that reasons through a harder request without stopping the conversation to do it — the difference between a voice agent that stalls while it works and one that keeps talking.
Both models handle 97 languages and, according to Google, detect a switch between them mid-conversation without being told. They also take visual input in near real time, so an agent can watch a camera feed while it listens, and they execute tool and API calls in the background rather than blocking on them.

The scores Google is quoting
Google places Extended Thinking first overall on Artificial Analysis’ Speech to Speech Quality Index with a score of 82.6. In the same announcement it reports 68.6% on the τ-Voice agentic benchmark, 35.1% on Sierra’s τ-Voice-banking variant, and 97.7% on Big Bench Audio. Gemini 3.8 Live, the cheaper model, ranks second on Speech Agent Arena.
Those are Google’s own numbers, taken from third-party leaderboards but published by the vendor on launch day. No independent re-run has appeared yet, and the leaderboards themselves had not refreshed to reflect the new models at the time of writing.
Priced under OpenAI’s voice model
Google’s developer post puts audio input at $0.005 a minute and audio output at $0.018 a minute. A minute of genuine two-way conversation therefore costs about $0.023 at list price.
That undercuts the most directly comparable product. OpenAI put its full-duplex GPT-Live-1 into the API at $0.05 a minute on 11 September. Voice has been the one modality where per-minute pricing is legible to buyers rather than buried in token maths, and Google has priced into that comparison deliberately.

A transcription model alongside it
Google also shipped Gemini 3.5 Transcribe, a speech-to-text model covering more than 85 languages. Google reports a word error rate of 4.0% in streaming mode and 2.6% when not streaming, with automatic code-switching and custom vocabulary biasing for up to 1,000 terms — the feature that matters for call centres full of product names a general model has never seen.
The Live models ship with integrations for Pipecat, LiveKit, LangChain, Agora, Fishjam and Vercel, which is a fair signal of who Google thinks the buyer is: teams already assembling voice agents from framework parts.
What to watch
The claims worth checking are the leaderboard ones. Artificial Analysis and Speech Agent Arena run their own evaluations on their own schedule, and whether Extended Thinking holds 82.6 and the top slot once they do is the test of Google’s launch-day framing. The second is latency under load: a model that reasons mid-sentence has to do it fast enough that the pause is not the product.