Gemini 3.5 Transcribe
Gemini 3.5 Transcribe is Google's speech-to-text model with 85+ languages, speaker diarization, word timestamps and about $0.005/min estimated cost.
What is this model and why does it matter?
Gemini 3.5 Transcribe is Google's dedicated speech-to-text model with 85+ language support, diarization, timestamps and smart transcription.
Gemini 3.5 Transcribe: features, use cases and important details
Gemini 3.5 Transcribe is Google’s dedicated non-streaming speech-to-text model, generally available since August 26, 2026.
Verified model facts
It supports automatic language detection across more than 85 languages/locales, speaker diarization, word-level timestamps, custom vocabulary and smart transcription that removes disfluencies and improves formatting.
Current status
It is active in the Gemini API with paid pricing based on audio input and text output tokens.
Best fit
It is best for meetings, interviews, podcasts, multilingual recordings and call-analysis pipelines.
Limitations
Some advanced features cannot be combined, diarization for three or more speakers is experimental, and it is not intended as a general audio-reasoning model.
Gemini 3.5 Transcribe capabilities and use cases
In addition, its main capabilities include 85+ languages, Automatic language detection, Speaker diarization, Word-level timestamps, Custom vocabulary and Smart transcription. For example, common use cases include Meetings, Podcasts, Interviews, Customer calls and Multilingual transcription.
Who should consider Gemini 3.5 Transcribe?
In practice, this model may suit Audio transcription, Meeting transcripts, Podcasts, Interviews, Call analytics and Multilingual speech-to-text. Also, notable strengths include Dedicated transcription model, Low estimated per-minute cost, Diarization and Word timestamps. However, review trade-offs such as Custom vocabulary cannot be combined with diarization/timestamps, Word timestamps may reduce accuracy and Designed for speech-to-text rather than audio reasoning before adopting it.
Gemini 3.5 Transcribe pricing and access
Meanwhile, Paid-tier pricing is $2.00 per 1M audio input tokens and $12.00 per 1M text output tokens, estimated at about $0.005 per transcribed minute blended. Google estimates an effective blended cost of about $0.005 per minute for standard non-streaming transcription.
Official resources and verification
Use the official model website, official documentation, pricing or release source and additional primary source to confirm current availability, limits and pricing. Product details can change after publication, so rely on primary documentation for final decisions.
Compare with other AI models
Next, continue your research in the AI models directory, Google models and Speech-to-Text Model models. Compare providers, pricing, modalities and practical limitations side by side to choose the right model for your workflow.
How to evaluate Gemini 3.5 Transcribe responsibly
First, test the model with a small set of realistic tasks before relying on it for production work. Also, check response quality, consistency, latency, supported file types, context limits and the effort required to review its output. For sensitive or regulated work, examine the provider’s privacy, data-retention, regional-processing and security documentation before submitting private information.
However, AI systems sometimes return incomplete, outdated or confidently incorrect information. Therefore, check important claims against trusted sources and test generated code before deployment. Pricing, quotas and model availability can also change without notice. Finally, revisit the official documentation before you plan a long-term integration or a large-volume workload.
How to use this model
- Create a Gemini API key.
- Upload an audio file.
- Call gemini-3.5-transcribe.
- Choose verbatim mode for timestamps/diarization or smart mode for cleaned-up text.
- Add custom vocabulary when domain terms need recognition biasing.
Example prompts
Transcribe this meeting and label each speaker.Return a verbatim transcript with word-level timestamps.Produce a cleaned smart transcript of this interview.
What it can do
- 85+ languages
- Automatic language detection
- Speaker diarization
- Word-level timestamps
- Custom vocabulary
- Smart transcription
Practical use cases
- Meetings
- Podcasts
- Interviews
- Customer calls
- Multilingual transcription
What does it cost?
Paid-tier pricing is $2.00 per 1M audio input tokens and $12.00 per 1M text output tokens, estimated at about $0.005 per transcribed minute blended.
What stands out
- Dedicated transcription model
- Low estimated per-minute cost
- Diarization
- Word timestamps
- Custom vocabulary
Things to consider
- Some feature combinations are incompatible
- Diarization beyond two speakers is experimental
- No general-purpose reasoning focus
Important restrictions and trade-offs
- Custom vocabulary cannot be combined with diarization/timestamps
- Word timestamps may reduce accuracy
- Designed for speech-to-text rather than audio reasoning
Our editorial take
A very competitive dedicated transcription API for teams that need multilingual speech recognition plus diarization and timestamps.