What Is Gemini Transcribe? 85+ Languages and Smart Formatting
Priya Nair
AI Infrastructure Analyst

TLDR85+ languages, live streaming, speaker labels, timestamps, and Smart formatting define Gemini Transcribe, Google's dedicated speech-to-text family.
Gemini Transcribe is Google's dedicated speech-to-text model family for converting recorded or live audio into text, with separate paths for complete recordings and streaming transcription. Google identifies the current models as gemini-3.5-transcribe and gemini-3.5-transcribe-live. The system supports Verbatim Mode, SMART Mode, Custom Vocabulary, and reported coverage of more than 85 languages. Early release coverage describes the service as a public preview, so pricing, quotas, and some limits remain subject to confirmation.
Key Takeaways
- Gemini Transcribe is specialized for speech recognition rather than general audio question-answering.
- The recorded-audio model uses the Interactions API, while Gemini Transcribe Live uses the Live API.
- SMART Mode removes fillers, repetitions, and false starts to produce cleaner text.
- Verbatim Mode preserves the spoken record and works with Speaker Diarization and Word-Level Timestamps on recorded audio.
- Automatic language detection is reported for more than 85 languages.
- A reported price of $5 per 1,000 minutes is unconfirmed and should not be treated as an official rate.
What Is Gemini Transcribe?
Gemini Transcribe is a Google speech-to-text model family built for developers who need a transcript as the primary output. Earlier Gemini models could analyze audio, answer questions about recordings, or produce summaries through prompts. Gemini Transcribe narrows that workflow around recognition, formatting, timestamps, speaker attribution, and live transcript delivery.
The family has two distinct interfaces. gemini-3.5-transcribe handles recorded audio through the Interactions API. gemini-3.5-transcribe-live handles continuous audio through the Live API, using a bidirectional streaming connection over WebSockets or the Google Gen AI SDK.
Google's Gemini API Cookbook lists Gemini Transcribe among its developer examples and describes support for audio files, live audio streams, word-level timestamps, speaker diarization, and custom vocabulary. The official live-transcription documentation identifies Live Transcription as a dedicated speech-recognition pipeline, separate from a conversational Live agent.
The release is not the same as the Gemini app's general ability to accept an audio file. It is an API-oriented transcription product intended for voice interfaces, captioning, dictation, media processing, and structured records. Early community coverage describes it as a public preview, while the supplied evidence does not establish a final production-status policy.
Gemini Transcribe at a Glance
| Specification | Current information |
|---|---|
| Developer | |
| Type | Dedicated speech-to-text model family |
| Model variants | gemini-3.5-transcribe and gemini-3.5-transcribe-live |
| Modality | Audio input to text output |
| Recorded-audio interface | Interactions API |
| Live interface | Live API over a streaming connection |
| Language coverage | More than 85 languages reported |
| Context window | Not yet confirmed |
| Pricing | Not yet confirmed; $5 per 1,000 minutes is an unconfirmed third-party report |
| Availability | Google developer APIs; public-preview status reported in early coverage |
| Downloadable weights | Not yet confirmed |
| License | Not yet confirmed |
| Live session limit | 10 minutes reported by a third-party release analysis; official confirmation not present in the bundle |
| Recorded-audio limit | Typically 1 hour reported by a third-party release analysis; official confirmation not present in the bundle |
The table separates documented product identity from claims that still need a Google pricing or limits page. That distinction matters for production planning. A model can be visible in developer documentation while its rate limits, retention terms, and regional availability remain incomplete in public references.
How Gemini Transcribe Works and What Makes It Different
Gemini Transcribe separates the transcription task from the reasoning task. A recorded file is submitted for transcription, and the service returns text with optional structure. A live audio stream produces partial text while a person speaks, followed by finalized text when the speech segment is complete.
Live Transcription exposes two result types. Interim transcription is a speculative, low-latency hypothesis for responsive captions. Finalized transcription represents the completed speech segment after a pause or turn boundary. Applications can render interim text immediately, then replace it with the finalized text.
Live input is described as raw 16-bit PCM audio. The live path is designed for continuous recognition rather than a turn-based assistant that listens, reasons, and speaks back. It can therefore act as the speech-to-text stage in a larger voice-agent pipeline.
The main product distinction is the choice between Verbatim Mode and SMART Mode:
- Verbatim Mode preserves fillers such as “um” and “uh,” repetitions, false starts, and self-corrections.
- SMART Mode removes unnecessary fillers, cleans up disfluencies, groups text into readable paragraphs or bullets, and formats dates, currencies, and numbers consistently.
- Custom Vocabulary lets developers bias recognition toward names, technical terms, and domain-specific jargon. The available descriptions report support for up to 1,000 terms.
- Speaker Diarization labels different speakers in recorded audio.
- Word-Level Timestamps attach start and end offsets to individual spoken words.
Gemini Transcribe's central design choice is to treat transcript cleanliness and evidentiary fidelity as separate output goals.
SMART Mode is useful when the transcript becomes a draft for reading. Verbatim Mode is better when the original wording matters. The available public descriptions say that recorded-audio SMART output cannot be combined with Speaker Diarization or Word-Level Timestamps. Live transcription does not return speaker labels or word-level timings in the current descriptions.
Early release material also reports automatic detection across more than 85 languages, along with language switching during a stream. The language count is broadly repeated, but detailed language coverage and switching behavior still require independent testing.
What You Can Do With Gemini Transcribe
Gemini Transcribe fits applications where audio must become searchable, editable, or actionable text.
Live captions and voice input: The Live API can stream interim and final transcript updates to a captioning interface. A 10-minute live-session ceiling is reported in early technical coverage, so longer events may require connection management or a different workflow.
Voice agents: Developers can place Gemini Transcribe Live at the front of a pipeline that sends text to an independent language model and then sends the response to a text-to-speech system. Third-party integrations have been reported for agent frameworks, but those reports are adoption signals rather than independent latency tests.
Meetings and interviews: The recorded-audio path can produce a Verbatim transcript with speaker labels and Word-Level Timestamps. That combination is suitable for review, quote selection, subtitle preparation, and records that need an audio trail.
Subtitles: Word-level timing can support subtitle-file generation and editing workflows. A community-built application using Gemini Flash demonstrates a related workflow with grouped timestamps and SRT export, but it is not Google's official product. The Gemini Transcribe community application should therefore be treated as an example implementation, not a reference implementation.
Clean dictation: SMART Mode can turn an unstructured voice memo into readable prose. It can also normalize numbers, dates, and currencies, reducing post-processing before the text enters a document or ticketing system.
Search and downstream extraction: Once audio is transcribed, a separate language model can summarize it, extract action items, or answer questions. Gemini Transcribe itself should not automatically be treated as the reasoning layer for those tasks.
How Gemini Transcribe Compares
The most relevant comparison is with a general Gemini model that understands audio. A general model can answer questions about a recording, while Gemini Transcribe is optimized for producing the transcript itself.
| Capability | Gemini Transcribe | General Gemini audio understanding |
|---|---|---|
| Primary task | High-throughput speech-to-text | Reasoning, analysis, and question-answering |
| Speaker labeling | Native support on recorded audio | Prompt-dependent |
| Timestamp precision | Word-level timestamps reported for recorded audio | Approximate timecodes may be prompt-generated |
| Cleanup control | Native Verbatim Mode and SMART Mode | Prompt-dependent |
| Live transcription | Dedicated Live API model | Not necessarily a dedicated transcription stream |
| Custom vocabulary | Native support reported for up to 1,000 terms | Prompt instructions alone may be less controlled |
For a separate look at the economics of a general Gemini tier, see What Gemini 3.7 Flash Actually Costs. That pricing analysis concerns Gemini 3.7 Flash, not Gemini Transcribe, and should not be used as a proxy for transcription pricing.
Based on the current public descriptions, Gemini Transcribe is the more natural fit when transcript structure, timing, or speaker identity is the deliverable. General Gemini audio models remain useful when the desired output is an interpretation of the recording.
Availability: How to Access Gemini Transcribe
The primary access route is Google's developer ecosystem. Google AI Studio is described as a place to test speech recognition, while the Gemini API provides programmatic access.
For recorded audio, the available implementation material describes uploading a file through Google's Files API and submitting its URI to the Interactions API with the gemini-3.5-transcribe model. For live audio, developers connect to the Live API and stream raw audio to gemini-3.5-transcribe-live.
The official documentation should be checked for current regions, quotas, supported SDK versions, authentication requirements, and retention terms before deployment. Those details are not fully present in the supplied evidence.
Kie.ai does not currently host Gemini Transcribe. Developers evaluating adjacent general-purpose Gemini workflows can also review Gemini 3.7 Flash, but it is a chat model rather than a dedicated speech-to-text endpoint.
What We Don't Know Yet
Several production-critical details remain open:
- The official price, including separate rates for recorded and live transcription, is not confirmed in the evidence bundle.
- A third-party analysis reports $5 per 1,000 minutes, but it does not establish an official Google rate.
- The live audio limit is reported as 10 minutes per session, while the recorded-audio limit is reported as typically 1 hour. Google confirmation is still needed.
- Speaker-count reporting is inconsistent. One third-party analysis reports up to 8 speakers, with 3 or more marked experimental, while early community posts cite a 3-speaker limit.
- The reported 2.6% to 5.5% word-error-rate range lacks dataset, language, noise, and baseline details.
- Exact support for language switching, regional availability, quotas, retention, and data-processing terms remains unconfirmed.
- Downloadable weights and an open-source license have not been established.
The reported WER range should be treated as a test hypothesis, not a purchasing specification. The most useful evaluation would test names, numbers, overlapping speech, background noise, speaker changes, and mixed-language audio on the recordings a production system actually receives.
Frequently Asked Questions
What is Gemini Transcribe?
Gemini Transcribe is Google's dedicated speech-to-text model family for converting recorded or live audio into text. It includes a recorded-audio model and a streaming Live API model.
Is Gemini Transcribe a model or an API?
Gemini Transcribe is a model family exposed through Google's APIs. The recorded-audio version is described through the Interactions API, while Gemini Transcribe Live runs through the Live API.
Does Gemini Transcribe work in real time?
Yes, Gemini Transcribe Live supports real-time transcription through a continuous Live API connection. It returns interim and finalized text as speech is received.
How many languages does Gemini Transcribe support?
Gemini Transcribe is reported to automatically detect more than 85 languages. The exact language-by-language support matrix is not yet confirmed in the supplied evidence.
Does Gemini Transcribe support speaker diarization?
The recorded-audio path supports speaker diarization, but the live path does not currently return speaker labels according to the available public descriptions. The maximum speaker count is not consistently confirmed.
How much does Gemini Transcribe cost?
Google's official price for Gemini Transcribe is not yet confirmed in the supplied evidence. A third-party release analysis reports $5 per 1,000 minutes, but that figure remains unconfirmed.
Is Gemini Transcribe open source?
Gemini Transcribe is not confirmed to be open source. The available material describes access through Google's hosted developer APIs rather than downloadable model weights.
What to watch next is Google's official pricing publication, confirmation of the 10-minute live and 1-hour recorded limits, and a reproducible benchmark covering the disputed WER and speaker-count claims.
Building similar multimodal AI workflows? On kie.ai you can try Gemini 3.8 Flash, Gemini 3.7 Flash, and Gemini 3.6 Flash.
About Priya Nair
Priya covers serving costs, context windows, and the infrastructure tradeoffs behind each model launch.
View all posts by Priya Nair