Google Launches Expressive Gemini 3.8 Flash TTS and Flash-Lite TTS Speech Models
Google has officially released Gemini 3.8 Flash TTS and Flash-Lite TTS, featuring over 100 languages, 2,000+ preset voices, line-by-line theatrical direction, a
On September 23, 2026, Google officially unveiled two new additions to its Gemini audio model family: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. Moving beyond static voice presets, these releases transform text-to-speech into a dynamic creative vocal studio capable of bespoke persona generation, prompt-driven direction, and expressive conversational speech.

Image source: Google
The launch cleanly divides responsibilities between two specialized engines: Flash TTS for deep creative direction and character design, and Flash-Lite TTS for high-volume, low-latency deployment.
Flash TTS vs. Flash-Lite TTS: Creative Direction Meets Real-Time Scale
Google structured the dual-model release to address distinct audio production workflows:
- Gemini 3.8 Flash TTS (Deep Creative Direction & Character Design): Tailored for high-fidelity production including gaming, immersive audiobooks, podcasts, and interactive media. It allows creators to build custom voice personas from scratch with natural language prompts, offering granular line-by-line control over acting cues, pacing, dialect shifts, and conversational backchanneling cues such as
<laughs>or active listening interjections like|mhm|. - Gemini 3.8 Flash-Lite TTS (Cost-Efficient Scale & Real-Time Agents): Built for scalable dubbing, bulk audio generation, and responsive voice agents. It automatically adapts tone and pacing on the fly, optimizing compute efficiency and latency for real-time customer and agent interactions.
Both models support long-form audio generation without auditory glitches, alongside dual-speaker screenplay staging from a single script with consistent voice retention and natural turn-taking.
2,000+ Production Voices, 30-Second Cloning, and Provenance Safeguards
Expanding beyond Google's earlier 30 voice presets, the models now provide a library of over 2,000 production-ready voices across more than 100 languages and dialects.
- Generative Persona Creation: Users can design completely new character voices or brand representatives using descriptive prompts and save them as reusable profiles.
- Flash TTS 30-Second Voice Replication: Flash TTS can replicate consistent vocal profiles from a 30-second audio sample, verified against a verbal consent recording from the voice owner.
- Safety and Watermarking: The voice replication workflow incorporates speaker-matched consent verification, while all synthesized audio carries Google's SynthID digital watermarking and C2PA provenance credentials.
Benchmarks and Platform Availability
According to Google, Gemini 3.8 Flash TTS and Flash-Lite TTS took the #1 and #2 spots respectively on Hume AI's Overall Quality Index. Google also stated that both models secured top positions in blind human preference evaluations on Voice Arena across key global languages, including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic (MSA), Mexican Spanish, and Hindi.
Both models begin rolling out immediately across Google's developer and consumer offerings:
- Developers: Available immediately in Google AI Studio and through the Gemini API.
- Consumers: Flash TTS is rolling out to Gemini Notebook (formerly NotebookLM), while Flash-Lite TTS powers audio generation in Google Vids.
- Enterprise: Both models are coming soon to Gemini Enterprise.