Google Releases Gemini 3.8 Flash TTS and Flash-Lite TTS for Expressive Voice Generation
Google has officially launched Gemini 3.8 Flash TTS and Flash-Lite TTS, bringing prompt-based voice design and Flash TTS 30-second voice cloning to Google AI St
On September 23, 2026, Google officially expanded its Gemini model family with the release of two next-generation audio generation models: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. The new speech models are immediately accessible in preview via Google AI Studio and the Gemini API, bringing expressive text-to-speech capabilities to developers, creators, and enterprise teams.

Image source: @GoogleAIStudio / Google
Extending Google's multimodal AI stack beyond text and vision, this release combines natural language voice design, Flash TTS's rapid 30-second voice replication, and nuanced accent modeling. The models are designed to power a wide range of audio applications, including conversational agents, interactive storytelling, video game dialogue, and automated podcast production.
Prompt-Based Voice Design and 30-Second Consent-Gated Voice Cloning
The centerpiece of the Gemini 3.8 TTS release is dynamic 'Voice Design', which enables creators to generate custom vocal styles, tones, and personas on the fly using natural language text prompts.
Developers can choose from an extensive library of more than 2,000 preset voices or define unique synthetic personas—such as a calm documentary narrator or an energetic game guide—entirely through prompt descriptions.
In addition, the flagship Gemini 3.8 Flash TTS model introduces rapid voice cloning from a 30-second reference audio sample. To mitigate potential misuse, deepfake exploitation, and unauthorized voice appropriation, Google has integrated mandatory consent verification checks directly into the cloning workflow, requiring verifiable authorization before reproducing any real-world voice profile.
Hume AI Benchmark Leadership and Flash vs. Flash-Lite Tier Differentiation
According to evaluation figures released by Google, Gemini 3.8 Flash TTS achieved first place overall on Hume AI's Voice Design Benchmark with a composite score of 71.4. In the benchmark's accent modeling evaluation, Flash TTS scored 60.8, leading the category in replicating authentic pronunciation and localized speech patterns.
On Hume AI's Overall Quality Index, Gemini 3.8 Flash TTS and the lightweight Gemini 3.8 Flash-Lite TTS secured first and second place respectively across evaluated models, demonstrating consistent synthesis fidelity across the product line.
The two model variants address distinct production requirements:
- Gemini 3.8 Flash TTS: The full-featured tier engineered for maximum emotional expressiveness, nuanced prosody, and 30-second voice cloning. It is targeted at rich narrative media, audiobooks, virtual avatars, and high-fidelity interactive storytelling.
- Gemini 3.8 Flash-Lite TTS: Built for high-volume, cost-efficient scale. It is optimized for high-volume dubbing, audio content creation, and expressive voice agents with fine-grained control over tone, pacing, and expressive nuance.
Multilingual Support and Dual-Speaker Screenplay Direction
The Gemini 3.8 TTS family supports multilingual generation across more than 100 languages. In addition to single-speaker synthesis, the API includes multi-speaker script directing capabilities, allowing developers to orchestrate multi-turn dialogue, tone shifts, and pacing between two distinct speakers within a single structured request.
Developers can immediately test prompt-based voice creation and fine-tune parameters using the Google AI Studio audio playground. For production pipelines, the models are integrated directly into the Gemini API (Interactions API / Speech Generation), enabling seamless programmatic orchestration within existing application workflows.
Because both models are currently rolling out in preview, regional availability and throughput quotas may vary depending on developer tier and infrastructure rollout stages. Teams planning production deployments should also factor in compliance workflows for the mandatory consent verification required when utilizing the 30-second voice cloning feature.