Gemini 2.5 Pro Preview TTS is Google’s premium text-to-speech model designed for studio-quality, high-fidelity audio generation. It enables developers to create realistic AI voices with natural language instructions, expressive speech, and long-form audio capabilities.

Pricing: Input 140 credits / 1M tokens (≈ $0.70), Audio Output 2,800 credits / 1M tokens (≈ $14.00). ~30% cheaper than official pricing. High-tier top-ups (+10% bonus) bring effective pricing down to ~90% of the above.
Input
0 / 2
No items yet. Click Add to start.

At least 1 item(s) required

0
No items yet. Click Add to start.

At least 1 item(s) required

Controls the randomness of the speech output. Higher values produce more creative and varied delivery, while lower values make the output more predictable and focused. Default value: 1

scene description

0/1000

Example Context / Overall Tone

0/1000
Output
output typeaudio

README

Complete guide to using google/gemini-2-5-pro-tts

Gemini 2.5 Pro Preview TTS API for Advanced AI Voice Generation

Original image
Studio-Quality AI Voice Generation with Gemini 2.5 Pro Preview TTS API
Expressive Voice and Style Control with Gemini 2.5 Pro Text-to-Speech API
Natural Multi-Speaker Dialogue Generation with Google Gemini 2.5 Pro TTS
Long-Form Audio Creation with Gemini 2.5 Pro Preview TTS API
FeatureGemini 2.5 Pro TTS APIGemini 2.5 Flash TTS API
Primary FocusPremium AI voice quality with advanced expression, natural speech, and professional audio generationFast and efficient AI voice generation optimized for low-latency applications
Best ForAudiobooks, podcasts, video narration, creative content, and long-form audio workflowsReal-time voice assistants, customer support agents, and high-volume voice applications
Voice QualityDelivers studio-quality speech with realistic pronunciation, natural prosody, emotional expression, and human-like deliveryProvides high-quality AI voices with faster generation speed for interactive experiences
Style & Emotion ControlOffers stronger style prompt following with detailed control over emotion, tone, pacing, and speaking styleSupports voice customization while focusing on speed and efficient generation
Multi-Speaker SupportDesigned for natural multi-speaker dialogue with consistent character voices and complex conversationsSupports conversational voice generation for interactive applications and voice agents
LatencyPrioritizes maximum audio quality and expressive performance over response speedOptimized for faster response times and real-time voice interactions
Long-Form ContentBetter suited for audiobooks, courses, documentaries, and extended narration projectsBetter suited for short-form content and applications requiring quick audio responses
Recommended Use CasesProfessional voice applications requiring premium quality, emotional depth, and natural storytellingAI assistants, scalable SaaS products, and applications where speed is the top priority