google / gemini-3-8-flash-lite-tts

Google’s efficient text-to-speech model built for low-latency, high-throughput voice generation. Gemini 3.8 Flash-Lite TTS supports natural speech across 101 languages, voice replication, multi-speaker audio, and long-form generation, making it ideal for voice agents and high-volume TTS workflows.

Pricing: Limited-time pricing (until 2026-12-31): Input 70 credits / 1M tokens (≈ $0.35), Audio Output 840 credits / 1M tokens (≈ $4.20). ~30% cheaper than official pricing. High-tier top-ups (+10% bonus) bring effective pricing down to ~90% of the above.
Input
0 / 2
No items yet. Click Add to start.

At least 1 item(s) required

0
No items yet. Click Add to start.

At least 1 item(s) required

When true, filler words like "hmm" and "ahh," interjections, and natural background sounds will be added to the conversation.

Controls the randomness of the speech output. Higher values produce more creative and varied delivery, while lower values make the output more predictable and focused. Default value: 1

Output
output typeaudio

README

Complete guide to using google/gemini-3-8-flash-lite-tts

Gemini 3.8 Flash-Lite TTS API – Try Google Gemini 3.8 TTS More Efficiently

Original image
Generate Speech with Low-Latency Gemini 3.8 Flash-Lite TTS
Scale High-Volume TTS with Better Cost Efficiency
Clone Voices from 30s Audio with Google Gemini 3.8 Flash-Lite TTS
Create Natural Multi-Speaker Conversations
Generate Long-Form Audio with Gemini 3.8 Text-to-Speech
Generate Speech in 101 Languages with Gemini 3.8 Flash-Lite TTS
Gemini 3.8 Flash-Lite TTS vs Flash TTS: What’s the Difference?
Access Both Gemini 3.8 TTS APIs on Kie AI