google / gemini-3-8-flash-tts

Google’s flagship creative text-to-speech model for high-fidelity, expressive voice generation. Gemini 3.8 Flash TTS supports natural-language voice design, voice cloning from around 30s of authorized audio, multi-speaker dialogue, and speech generation across 130 languages.

Pricing: Limited-time pricing (until 2026-12-31): Input 70 credits / 1M tokens (≈ $0.35), Audio Output 1,260 credits / 1M tokens (≈ $6.30). ~30% cheaper than official pricing. High-tier top-ups (+10% bonus) bring effective pricing down to ~90% of the above.
Input
0 / 2
No items yet. Click Add to start.

At least 1 item(s) required

0
No items yet. Click Add to start.

At least 1 item(s) required

When true, filler words like "hmm" and "ahh," interjections, and natural background sounds will be added to the conversation.

Controls the randomness of the speech output. Higher values produce more creative and varied delivery, while lower values make the output more predictable and focused. Default value: 1

Output
output typeaudio

README

Complete guide to using google/gemini-3-8-flash-tts

Gemini 3.8 Flash TTS API – Create Custom Voices with 30-Second Voice Cloning

Original image
Design Original Voices from Natural Language
Replicate a Voice from Just 30 Seconds of Audio
Gemini 3.8 Flash TTSGemini 3.8 Flash-Lite TTS
Primary FocusVoice fidelity & creative controlSpeed & cost efficiency
Supported Languages130101
Best ForCreative & professional audioHigh-volume speech generation
Long-Form NarrationSupportedSupported
Voice ReplicationSupportedSupported
Typical WorkloadsAudiobooks, dialogue, charactersVoice agents, read-aloud, bulk audio