Google released two text-to-speech models on 23 September, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, and the headline number is the voice library: more than 2,000 production-ready voices, against the 30 the previous Gemini speech models offered, according to Unite.AI’s account. Coverage runs to more than 100 languages and dialects, including regional varieties Google names specifically: Mexican Spanish, Quebec French, Scots English.
Both models are available now through the Gemini API and Google AI Studio.
Two models, two jobs
Google describes Flash TTS as built for deep creative direction and character design — games, audiobooks, podcasts — and Flash-Lite TTS as the high-volume, cost-efficient option, aimed at dubbing and voice agents. Google did not publish pricing for either in the announcement post.

The capability that matters more than the voice count is voice design. Rather than picking from a list, a developer describes the voice they want in ordinary language — the role, the accent, the character of the delivery — and the model generates it. Google also supports line-by-line performance direction, where script cues steer the read, native two-speaker staging for conversations, and long-form generation that it says holds up across hours of audio.
Cloning, consent and the map
The models will also replicate a voice from a 30-second audio sample, which Google gates behind a consent verification step. Output carries a SynthID watermark.
That combination is where the release stops being a product update. Voice replication through AI Studio is unavailable in Illinois, Texas, the European Economic Area, the United Kingdom, Switzerland and India — a list that reads as a map of biometric and AI law rather than a list of markets. Illinois and Texas both have biometric identifier statutes; the EEA and the UK have data protection regimes that treat a voiceprint as sensitive.

The benchmark claim
Google reports that Flash TTS took first place on Hume AI’s Voice Design Benchmark with 71.4, and first in accent modelling with 60.8. Hume AI is a third party, but the scores as presented are Google’s citation of them in its own launch post, not an independently run evaluation of this release. Unite.AI reports that Flash TTS and Flash-Lite TTS placed first and second on Hume’s overall quality index.
What to watch
Two things. Whether Google publishes per-character or per-second pricing for Flash-Lite, since dubbing economics turn entirely on that. And whether the consent-verification step for voice replication proves strong enough to keep the feature available in the jurisdictions currently excluded — or whether that exclusion list grows.