Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Speech to Text Model

Scribe v2

Scribe v2 is a speech-to-text transcription model from ElevenLabs that supports speaker identification.

PublisherElevenLabs
TypeTranscription
Released2026
Price$0.0067/min
LATEST

Automatic speech recognition with speaker diarization

Scribe v2 is a speech-to-text transcription model developed by ElevenLabs, the audio AI company known for its voice synthesis products. It converts spoken audio into written text and includes an option to identify and label individual speakers within a recording, a feature known as speaker diarization. The model was added to MindStudio in March 2026 and carries the LATEST tag, indicating it is ElevenLabs' current generation transcription model.

Scribe v2 is suited for workflows that require converting audio content into structured text, such as transcribing interviews, meetings, podcasts, or voice recordings where distinguishing between multiple speakers is useful. The speaker identification feature is exposed as a configurable input, allowing users to toggle it on or off depending on their use case. It operates as a dedicated transcription model rather than a general-purpose language model, meaning it does not generate text responses or perform reasoning tasks.

What Scribe v2 supports

Audio Transcription

Converts spoken audio input into written text. Designed specifically for transcription tasks rather than generative language output.

Speaker Diarization

Identifies and labels individual speakers within an audio recording. Toggled via the 'Include Speakers' input option.

Latest Model Version

Tagged as the current LATEST release in ElevenLabs' transcription model lineup as of March 2026.

Configurable Output

Exposes a select-type input for speaker inclusion, allowing users to adjust transcription behavior without code changes.

Ready to build with Scribe v2?

Get Started Free

Common questions about Scribe v2

Does Scribe v2 have a context window limit?

No context window size is specified in the available metadata for Scribe v2. As a transcription model, it processes audio input rather than text tokens, so a traditional context window measurement may not apply.

What is the pricing for Scribe v2?

Pricing information for Scribe v2 is not listed in the available metadata. You should check ElevenLabs' official pricing page or MindStudio's billing documentation for current rates.

Does Scribe v2 support identifying multiple speakers?

Yes. Scribe v2 includes a speaker diarization option exposed as the 'Include Speakers' input. When enabled, the model labels different speakers within the transcribed output.

What types of tasks is Scribe v2 designed for?

Scribe v2 is a dedicated speech-to-text model. It is intended for transcription tasks such as converting interviews, meetings, podcasts, or voice recordings into text, and is not designed for text generation or reasoning.

When was Scribe v2 released?

Scribe v2 was added to MindStudio on March 13, 2026, and carries a release date of 2026. It is tagged as the LATEST version of ElevenLabs' transcription model.

Parameters & options

Include SpeakersSelect

Choose whether to include timing and speaker information in the transcription

Default: no
YesNo
Include Word TimingsSelect

Add per-word start/end times and confidence inside each transcript segment. Increases the size of the result considerably.

Default: no
YesNo

Start building with Scribe v2

No API keys required. Create AI-powered workflows with Scribe v2 in minutes — free.