Google introduces Gemini 3.8 Flash speech generation models
Google introduced Gemini 3.8 Flash TTS and Flash-Lite TTS, two speech generation models supporting more than 100 languages. Flash TTS allows users to create new voices from text descriptions or use a library of more than 2,000 preset voices.
The models support voice cloning from a 30-second audio sample, stage directions for dialogue, and an inaudible SynthID watermark. Both models are rolling out through the Gemini API, Google AI Studio, Gemini Notebook, and Google Vids, with Gemini Enterprise API access following soon.