Sponsored by Byond Boundrys Consulting - Empowering Ideas, Delivering Results
ElevenLabs Beginner-friendly

Scribe v2

Scribe v2 is a verified current AI model with official specifications, pricing or access information, capabilities, practical use cases and limitations.

Speech Recognition ModelTextAudioVideo Freemium
In plain English

What is this model and why does it matter?

Scribe v2 is ElevenLabs' state-of-the-art speech recognition model with 90+ language support, word-level timestamps, speaker diarization, keyterm prompting, entity detection and audio tagging.

Meeting transcriptionCall analyticsPodcastsCaptionsMultilingual speech-to-textMedia indexing
Model overview

Scribe v2: features, use cases and important details

Scribe v2 is a current AI model verified from first-party ElevenLabs sources.

Scribe v2 verified specifications

Scribe v2 is ElevenLabs’ state-of-the-art speech recognition model with 90+ language support, word-level timestamps, speaker diarization, keyterm prompting, entity detection and audio tagging. The verified context or usage limit is Audio transcription workflow; files up to 3 GB, with Text transcript with timestamps and metadata maximum output.

Scribe v2 pricing and access

ElevenLabs lists Scribe v2 API speech-to-text pricing at $0.22 per hour. Optional keyterm prompting and entity detection can add usage cost.

Scribe v2 best uses

Meeting transcription, Call analytics, Podcasts, Captions, Multilingual speech-to-text, Media indexing.

Scribe v2 limitations

Accuracy depends on audio quality, Specialized terms may need keyterms, Realtime use requires the separate Scribe v2 Realtime model.

Get started

How to use this model

  1. Create an ElevenLabs API key.
  2. Upload audio or video.
  3. Select model scribe_v2.
  4. Enable timestamps, speaker diarization or keyterms as needed.
  5. Review the transcript and detected entities.
Copy and try

Example prompts

  • Transcribe this meeting with speaker labels.
  • Generate word-level timestamps for this interview.
  • Use these product names as keyterms while transcribing.
Capabilities

What it can do

  • 90+ languages
  • Word timestamps
  • Up to 32 speakers
  • Keyterm prompting
  • Entity detection
  • Audio tagging
Best for

Practical use cases

  • Meetings
  • Contact centers
  • Media transcription
  • Subtitles
  • Searchable audio archives
Pricing

What does it cost?

ElevenLabs lists Scribe v2 API speech-to-text pricing at $0.22 per hour. Optional keyterm prompting and entity detection can add usage cost.

Input$0.22 / audio hour
OutputIncluded in transcription pricing
Simple summaryBase API transcription is $0.22 per audio hour before optional feature surcharges.

What stands out

  • Low hourly price
  • High language coverage
  • Rich transcription metadata
  • Speaker diarization

Things to consider

  • Speech-to-text only
  • Optional features may cost more
  • Proprietary
Limitations

Important restrictions and trade-offs

  • Accuracy depends on audio quality
  • Specialized terms may need keyterms
  • Realtime use requires the separate Scribe v2 Realtime model
SimplifyAITools verdict

Our editorial take

A highly practical current transcription model for teams that need accurate multilingual speech-to-text with timestamps, speaker labels and enterprise-oriented extraction features.

References

Primary sources

  1. Open source 1 ↗
  2. Open source 2 ↗
  3. Open source 3 ↗