Speech generation
Create rich custom voices, replicate reference audio, and orchestrate two-speaker dialogue with control over style, pitch and performance.
Create rich custom voices, replicate reference audio, and orchestrate two-speaker dialogue with control over style, pitch and performance.
Generate unique life-like vocal identities — or tap into over a thousand production-ready voices in our pre-crafted library.
Best for generative voice design and two-speaker dialogue generation. Keeps intonation, inflection, and pace steady from sentence-to-sentence, as you direct with precision.
Best for everyday audio generation, like video dubbing or podcast creation. Balances efficiency with expressiveness to power large content libraries and high-traffic apps.
Generate speech with precise control to create emotive narratives – or create entire conversations featuring multiple speakers.
Create a rich cast of voices using simple natural language prompts. Describe any persona — from warm documentary narrator to eccentric fantasy character — or fine-tune traits like accent, age, pitch, and texture on demand.*
*Gemini 3.8 Flash TTS only.
Direct single-narrator and multi-character scenes with clear voice distinction. Separate dialogue from acting notes for control over emotion, tone, and volume.
Choose from thousands of production-ready voices in different global accents, dialects and languages.
Turn seconds of reference audio into a high-fidelity digital voice. Our secure, consent-based workflow lets you create accurate voice copies for dubbing, narration, and scaling voice talent.
Use intuitive audio tags to command style, pace, and delivery with unprecedented precision.
Our speech generation models deliver impressively fast speech generation without compromising on vocal stability or expressive quality.
On text-to-speech evals, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS secure the #1 and #2 spots on Hume AI’s Overall Quality Index, driving major improvements in long-form stability and dual-speaker screenplay control compared to Gemini 3.1 Flash TTS.
| Benchmark | Gemini 3.8 Flash TTS | Gemini 3.8 Flash-Lite TTS | Gemini 3.1 Flash TTS | Eleven Labs v3 | Eleven Labs v3 conversational | Cartesia Sonic 3.6 | OpenAI gpt-4o-mini-tts | Inworld TTS-2 |
|---|---|---|---|---|---|---|---|---|
| Overall Reliability x Expressiveness | 0.920 | 0.914 | 0.783 | 0.706 | 0.769 | 0.840 | 0.740 | 0.576 |
| Human-like variation Lower = flatter or wilder than a human across turns | 4.58 | 4.51 | 3.95 | 5.00 | 4.97 | 3.40 | 4.22 | 3.76 |
| Multispeaker | 4.14 | 4.10 | 3.60 | 3.85 | — | — | — | — |
| Style Tag Control Single Tag | 4.34 | 4.32 | 4.31 | 3.89 | 3.97 | 3.37 | — | 4.15 |
Gemini 3.8 Flash TTS’s voice design capabilities deliver frontier-level customization, securing the #1 overall spot on Hume AI’s Voice Design Benchmark (71.4) and leading the industry in accent modeling (60.8).
| Benchmark | Gemini 3.8 Flash TTS | ElevenLabs Voice Design v3 | Inworld Voice Design |
|---|---|---|---|
| Overall English | 71.4 | 70.8 | 69.8 |
| Multilingual | 3.82 | 3.65 | 3.57 |
| Accents | 60.8 | 45.4 | 35.8 |
| Voice Qualities Single Tag | 74.6 | 76.6 | 76.3 |
In blind human preference evaluations on Voice Arena, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS secure top positions amongst competitors in key global languages, including English, Japanese, Brazilian Portuguese, Vietnamese, and Modern Standard Arabic (MSA), Mexican Spanish and Hindi. With support for over 100 languages, these models empower creators, developers, and enterprises to build high-quality, multilingual voice experiences worldwide.
| Benchmark | Gemini 3.8 Flash TTS | Gemini 3.8 Flash-Lite TTS | Gemini 3.1 Flash TTS | ElevenLabs v3 | Cartesia Sonic 3.6 | OpenAI gpt-4o-mini-tts |
|---|---|---|---|---|---|---|
| English | 1061 | 1087 | 1051 | 986 | 1068 | 940 |
| Japanese | 1232 | 1152 | 1148 | 1048 | — | 975 |
| Brazilian Portuguese | 1104 | 1134 | 1094 | 1040 | 1080 | 946 |
| Vietnamese | 1135 | 1156 | 1099 | 1043 | — | 839 |
| Arabic MSA | 1204 | 1181 | 1135 | 1020 | — | 911 |
| Hindi | 1106 | 1076 | 1086 | 1052 | 1104 | 843 |
| Mexican Spanish | 1152 | 1146 | 1092 | 1015 | 1089 | 880 |
| Name | 3.8 Flash TTS | 3.8 Flash-Lite TTS |
|---|---|---|
| Status | General availability | General availability |
| Input |
|
|
| Output |
|
|
| Input tokens | 8K | 8K |
| Output tokens | 64K | 64K |
| Availability |
|
|
| Documentation | View developer docs | View developer docs |
| Model card | View model card | View model card |
The fastest path from prompt to production
AI-powered video creation for work
Get started building with cutting-edge AI models
Your research and thinking partner