Google's Gemini learned to speak in your voice from a 30-second clip.

Google released two speech models on 23 September, Gemini 3.8 Flash TTS and Flash-Lite TTS, in its tools for developers. Google says they can "replicate" a voice from "a 30-second audio sample", with "built-in consent verification". In Google's film, the user records a sample and then reaches a step that reads "Verify it's you".
Developers can also describe a voice in words ("Woman in her 30's, pumped, emphatic with Brooklyn accent" is the film's example) or pick one of "2,000+ production-ready voices" in more than 100 languages and dialects, Google says. Every clip carries SynthID, Google's inaudible watermark.
Voice cloning through AI Studio, Google's developer site, is not available in the EEA, the UK, Switzerland, India, Illinois or Texas, Google's blog says. It gives no reason. Support in Gemini Enterprise, the version sold to companies, is "coming soon".
Google's film carries its own note: sequences shortened and screen images simulated.
Would your company let an employee's cloned voice answer its help line?
🎥 Video: Google DeepMind, official YouTube channel · narration: AI voice
What the captions say
- Google's Gemini can now copy a voice.
- It needs 30 seconds of audio, plus consent.
- Or you type the voice you want.
- Released on 23 September for developers.
- Over 2,000 voices, 100+ languages.
- Cloning is not offered in the EEA, UK or Switzerland.
- Google's film: screens simulated, sequences shortened.


