Gemini Audio

Speech generation

Create rich custom voices, replicate reference audio, and orchestrate two-speaker dialogue with control over style, pitch and performance.

Models

Generate unique life-like vocal identities — or tap into over a thousand production-ready voices in our pre-crafted library.


Capabilities

Generate speech with precise control to create emotive narratives – or create entire conversations featuring multiple speakers.

Generative voice design

Create a rich cast of voices using simple natural language prompts. Describe any persona — from warm documentary narrator to eccentric fantasy character — or fine-tune traits like accent, age, pitch, and texture on demand.*

*Gemini 3.8 Flash TTS only.

Scripted dialogue control

Direct single-narrator and multi-character scenes with clear voice distinction. Separate dialogue from acting notes for control over emotion, tone, and volume.

Voice library and regional presets

Choose from thousands of production-ready voices in different global accents, dialects and languages.

Voice replication

Turn seconds of reference audio into a high-fidelity digital voice. Our secure, consent-based workflow lets you create accurate voice copies for dubbing, narration, and scaling voice talent.

Granular expressive control

Use intuitive audio tags to command style, pace, and delivery with unprecedented precision.


Performance

Our speech generation models deliver impressively fast speech generation without compromising on vocal stability or expressive quality.

Text-to-Speech Quality Benchmark
Hume AI

On text-to-speech evals, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS secure the #1 and #2 spots on Hume AI’s Overall Quality Index, driving major improvements in long-form stability and dual-speaker screenplay control compared to Gemini 3.1 Flash TTS.

BenchmarkGemini 3.8 Flash TTSGemini 3.8 Flash-Lite TTSGemini 3.1 Flash TTSEleven
Labs v3Eleven
Labs v3 conversationalCartesia Sonic 3.6OpenAI gpt-4o-mini-ttsInworld TTS-2
Overall Reliability x Expressiveness0.9200.9140.7830.7060.7690.8400.7400.576
Human-like variation Lower = flatter or wilder
than a human across turns4.584.513.955.004.973.404.223.76
Multispeaker4.144.103.603.85————
Style Tag Control Single Tag4.344.324.313.893.973.37—4.15

Text-to-Speech Voice Design Leaderboard
Hume AI

Gemini 3.8 Flash TTS’s voice design capabilities deliver frontier-level customization, securing the #1 overall spot on Hume AI’s Voice Design Benchmark (71.4) and leading the industry in accent modeling (60.8).

BenchmarkGemini 3.8 Flash TTSElevenLabs Voice Design v3Inworld Voice Design
Overall English71.470.869.8
Multilingual3.823.653.57
Accents60.845.435.8
Voice Qualities Single Tag74.676.676.3

Text-to-Speech Leaderboard
Voice Arena

In blind human preference evaluations on Voice Arena, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS secure top positions amongst competitors in key global languages, including English, Japanese, Brazilian Portuguese, Vietnamese, and Modern Standard Arabic (MSA), Mexican Spanish and Hindi. With support for over 100 languages, these models empower creators, developers, and enterprises to build high-quality, multilingual voice experiences worldwide.

BenchmarkGemini 3.8 Flash TTSGemini 3.8 Flash-Lite TTSGemini 3.1 Flash TTSElevenLabs v3Cartesia Sonic 3.6OpenAI gpt-4o-mini-tts
English1061108710519861068940
Japanese1232115211481048—975
Brazilian Portuguese11041134109410401080946
Vietnamese1135115610991043—839
Arabic MSA1204118111351020—911
Hindi11061076108610521104843
Mexican Spanish11521146109210151089880

Model information

Name3.8 Flash TTS3.8 Flash-Lite TTS
StatusGeneral availabilityGeneral availability
Input
  • Text
  • Text
Output
  • Audio
  • Audio
Input tokens8K8K
Output tokens64K64K
Availability
  • Google AI Studio
  • Gemini API
  • Gemini Notebook
  • Google AI Studio
  • Gemini API
  • Google Vids
DocumentationView developer docsView developer docs
Model cardView model cardView model card

Try Speech generation