What was added
Google has added real-time video to its live conversational model. Gemini 3.8 Live with Live Avatar, announced on Thursday, generates a speaking face while the model talks, with lip-sync, facial expression and turn-taking handled as part of the same stream rather than bolted on afterwards.
The language claim is the striking part. Google says the system handles 97 languages with native speech-to-speech synchronisation, adapting lip movement and expression when a conversation switches language mid-way. It also supports asynchronous tool calls, so the avatar can keep talking while the system fetches something in the background — the awkward silence in an agent call is a product problem as much as a technical one.
The release lands a day after Google shipped new Gemini 3.8 text-to-speech models with more than 2,000 voices, so the voice layer and the face layer arrived within twenty-four hours of each other. Taken together they describe a product direction rather than two features: Google is assembling an assistant that can be deployed as a presence, not as a text box.
Where it runs, and who can change the face
This is an enterprise release rather than a consumer one. Live Avatar is generally available in Gemini Enterprise, with US and EU endpoints, provisioned throughput and the compliance apparatus that implies. Google did not publish pricing with the announcement.

Customers can pick from preset avatars, and companies on an allowlist can generate a custom one from a reference image. That second option is the one with consequences: an allowlist is a policy, not a technical limit, and it is the only thing standing between this and a generated face that looks like a specific person.
The watermark, and what it does not do
Google says all output is watermarked with SynthID, in both audio and video. That is a real measure and a partial one. SynthID is designed to survive ordinary handling and to be detectable by Google’s own tooling; it is not a consent mechanism, and it does nothing about a video that has been re-recorded off a screen.

What to watch
Three things will show whether this is a product or a demo. Whether Google publishes per-minute pricing, since real-time video generation is expensive to serve. Whether the allowlist for custom avatars has stated criteria, or is granted case by case. And whether any of the 97 languages arrive with quality figures attached, because a lip-sync that works in English and slips in Hindi is a different product in each market.