Scribe v1
Scribe v1 is a speech-to-text transcription model from ElevenLabs, available on MindStudio for audio transcription tasks.
Automatic speech recognition from ElevenLabs
Scribe v1 is a speech-to-text model developed by ElevenLabs, released in July 2025. It is designed to convert spoken audio into written text and supports speaker identification through its "Include Speakers" input option, which allows transcripts to be labeled by individual speakers. ElevenLabs is known primarily for its voice synthesis technology, and Scribe v1 represents the company's entry into the transcription space as a first-party offering.
Scribe v1 is suited for workflows that require accurate audio transcription, particularly in scenarios involving multiple speakers such as interviews, meetings, or podcasts. The model is available directly on MindStudio without requiring separate API key configuration. Because it is a transcription model rather than a language model, it does not have a context window in the traditional sense, and its output is determined by the length and content of the audio input provided.
What Scribe v1 supports
Audio Transcription
Converts spoken audio input into written text. Designed for transcription tasks across a range of audio formats and recording conditions.
Speaker Diarization
Identifies and labels individual speakers within a transcript using the "Include Speakers" input option. Useful for multi-speaker audio such as interviews or meetings.
Structured Transcript Output
Returns transcription results in a structured format that can be used downstream in MindStudio workflows. Output is segmented to reflect the audio content.
Language Recognition
Detects and transcribes spoken language from audio input. Scribe v1 is built to handle natural speech patterns across different speakers and recording environments.
Ready to build with Scribe v1?
Get Started FreeCommon questions about Scribe v1
Does Scribe v1 have a context window?
No. Scribe v1 is a transcription model, not a language model, so it does not have a context window. Its output is determined by the duration and content of the audio file provided.
What does the 'Include Speakers' option do?
The 'Include Speakers' input is a select-type option that, when enabled, adds speaker labels to the transcript output. This is useful for audio with multiple participants, such as interviews or panel discussions.
What types of audio can Scribe v1 transcribe?
Scribe v1 is designed for general speech-to-text transcription. While the metadata does not specify supported file formats explicitly, ElevenLabs' Scribe model is publicly documented to support common audio formats.
Is there a published price for using Scribe v1 on MindStudio?
No published price is listed in the available metadata for Scribe v1. Pricing details would be available through MindStudio's platform or ElevenLabs' official documentation.
When was Scribe v1 released?
Scribe v1 was released in July 2025 and was added to the MindStudio model catalog on August 8, 2025.
What people think about Scribe v1
Community discussions referencing Scribe v1 appear primarily in the context of broader speech-to-text benchmarking threads on r/LocalLLaMA. Participants in these threads evaluated ElevenLabs' transcription models alongside other cloud and local STT options, particularly for long-form and medical dialogue use cases.
A recurring theme in these discussions is the use of medical dialogue as a benchmark domain, where transcription accuracy and handling of specialized terminology are key concerns. The threads do not focus exclusively on Scribe v1 and cover a wide range of competing models, suggesting users evaluate it as one option among many rather than a standalone solution.
I benchmarked 26 local + cloud Speech-to-Text models on long-form medical dialogue and ranked them + open-sourced the full eval
Benchmark: 15 STT models on long-form medical dialogue
Documentation & links
Parameters & options
Choose whether to include timing and speaker information in the transcription
Add per-word start/end times and confidence inside each transcript segment. Increases the size of the result considerably.
Explore similar models
Start building with Scribe v1
No API keys required. Create AI-powered workflows with Scribe v1 in minutes — free.