Google's new Flash TTS models let you design AI voices from scratch using text descriptions

Google is introducing two new text-to-speech models, Gemini 3.8 Flash TTS and Flash-Lite TTS, which support more than 100 languages. Flash TTS can create new voices from text descriptions, and both models let users add stage directions to individual lines and generate two-voice dialogue from a single script. A voice cloning feature can build a voice profile from a 30-second sample. The article…
This is a summary curated by AIFuture. Read the complete article at the original source:
Read the full story on The Decoder