Sponsored by Byond Boundrys Consulting - Empowering Ideas, Delivering Results
Google New Intermediate

Gemini 3.5 Transcribe

Gemini 3.5 Transcribe is Google's speech-to-text model with 85+ languages, speaker diarization, word timestamps and about $0.005/min estimated cost.

Speech-to-Text ModelTextAudio Paid
In plain English

What is this model and why does it matter?

Gemini 3.5 Transcribe is Google's dedicated speech-to-text model with 85+ language support, diarization, timestamps and smart transcription.

Audio transcriptionMeeting transcriptsPodcastsInterviewsCall analyticsMultilingual speech-to-text
Model overview

Gemini 3.5 Transcribe: features, use cases and important details

Gemini 3.5 Transcribe is Google’s dedicated non-streaming speech-to-text model, generally available since August 26, 2026.

Verified model facts

It supports automatic language detection across more than 85 languages/locales, speaker diarization, word-level timestamps, custom vocabulary and smart transcription that removes disfluencies and improves formatting.

Current status

It is active in the Gemini API with paid pricing based on audio input and text output tokens.

Best fit

It is best for meetings, interviews, podcasts, multilingual recordings and call-analysis pipelines.

Limitations

Some advanced features cannot be combined, diarization for three or more speakers is experimental, and it is not intended as a general audio-reasoning model.

Gemini 3.5 Transcribe capabilities and use cases

In addition, its main capabilities include 85+ languages, Automatic language detection, Speaker diarization, Word-level timestamps, Custom vocabulary and Smart transcription. For example, common use cases include Meetings, Podcasts, Interviews, Customer calls and Multilingual transcription.

Who should consider Gemini 3.5 Transcribe?

In practice, this model may suit Audio transcription, Meeting transcripts, Podcasts, Interviews, Call analytics and Multilingual speech-to-text. Also, notable strengths include Dedicated transcription model, Low estimated per-minute cost, Diarization and Word timestamps. However, review trade-offs such as Custom vocabulary cannot be combined with diarization/timestamps, Word timestamps may reduce accuracy and Designed for speech-to-text rather than audio reasoning before adopting it.

Gemini 3.5 Transcribe pricing and access

Meanwhile, Paid-tier pricing is $2.00 per 1M audio input tokens and $12.00 per 1M text output tokens, estimated at about $0.005 per transcribed minute blended. Google estimates an effective blended cost of about $0.005 per minute for standard non-streaming transcription.

Official resources and verification

Use the official model website, official documentation, pricing or release source and additional primary source to confirm current availability, limits and pricing. Product details can change after publication, so rely on primary documentation for final decisions.

Compare with other AI models

Next, continue your research in the AI models directory, Google models and Speech-to-Text Model models. Compare providers, pricing, modalities and practical limitations side by side to choose the right model for your workflow.

How to evaluate Gemini 3.5 Transcribe responsibly

First, test the model with a small set of realistic tasks before relying on it for production work. Also, check response quality, consistency, latency, supported file types, context limits and the effort required to review its output. For sensitive or regulated work, examine the provider’s privacy, data-retention, regional-processing and security documentation before submitting private information.

However, AI systems sometimes return incomplete, outdated or confidently incorrect information. Therefore, check important claims against trusted sources and test generated code before deployment. Pricing, quotas and model availability can also change without notice. Finally, revisit the official documentation before you plan a long-term integration or a large-volume workload.

Get started

How to use this model

  1. Create a Gemini API key.
  2. Upload an audio file.
  3. Call gemini-3.5-transcribe.
  4. Choose verbatim mode for timestamps/diarization or smart mode for cleaned-up text.
  5. Add custom vocabulary when domain terms need recognition biasing.
Copy and try

Example prompts

  • Transcribe this meeting and label each speaker.
  • Return a verbatim transcript with word-level timestamps.
  • Produce a cleaned smart transcript of this interview.
Capabilities

What it can do

  • 85+ languages
  • Automatic language detection
  • Speaker diarization
  • Word-level timestamps
  • Custom vocabulary
  • Smart transcription
Best for

Practical use cases

  • Meetings
  • Podcasts
  • Interviews
  • Customer calls
  • Multilingual transcription
Pricing

What does it cost?

Paid-tier pricing is $2.00 per 1M audio input tokens and $12.00 per 1M text output tokens, estimated at about $0.005 per transcribed minute blended.

Input$2.00 / 1M audio tokens
Output$12.00 / 1M text tokens
Simple summaryGoogle estimates an effective blended cost of about $0.005 per minute for standard non-streaming transcription.

What stands out

  • Dedicated transcription model
  • Low estimated per-minute cost
  • Diarization
  • Word timestamps
  • Custom vocabulary

Things to consider

  • Some feature combinations are incompatible
  • Diarization beyond two speakers is experimental
  • No general-purpose reasoning focus
Limitations

Important restrictions and trade-offs

  • Custom vocabulary cannot be combined with diarization/timestamps
  • Word timestamps may reduce accuracy
  • Designed for speech-to-text rather than audio reasoning
SimplifyAITools verdict

Our editorial take

A very competitive dedicated transcription API for teams that need multilingual speech recognition plus diarization and timestamps.

References

Primary sources

  1. Open source 1 ↗
  2. Open source 2 ↗
  3. Open source 3 ↗