Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Speech to Text Model

Whisper Large v3

Whisper Large v3 is an open-source speech-to-text model from OpenAI that transcribes and translates audio across 99 languages.

PublisherOpenAI
TypeTranscription
ReleasedNovember 2023
Input$7.50/MTok
OutputFree/MTok
ProviderDeepInfra
OPEN SOURCELOW COST

Multilingual speech recognition and transcription

Whisper Large v3 is an automatic speech recognition model developed by OpenAI and released in November 2023. It is trained on a large dataset of multilingual and multitask supervised data collected from the web, enabling it to transcribe audio in 99 languages and translate speech into English. The model uses a Transformer-based encoder-decoder architecture and is available as open source under the MIT license.

Whisper Large v3 is well suited for tasks that require accurate transcription of spoken audio, including meeting notes, subtitles, voice interfaces, and multilingual content processing. It is hosted on DeepInfra, making it accessible via API without requiring users to manage their own infrastructure. The model is tagged as low cost, making it a practical option for high-volume transcription workloads.

What Whisper Large v3 supports

Speech Transcription

Converts spoken audio into written text across 99 languages using a Transformer encoder-decoder architecture.

Speech Translation

Translates spoken audio from supported languages directly into English text in a single pass.

Open Source Access

Released under the MIT license, allowing free use, modification, and redistribution of model weights and code.

Low Cost Inference

Hosted on DeepInfra at low per-request pricing, making it practical for high-volume audio transcription workloads.

Multilingual Support

Handles audio input in 99 languages, trained on a diverse multilingual dataset sourced from the web.

Ready to build with Whisper Large v3?

Get Started Free

Common questions about Whisper Large v3

How many languages does Whisper Large v3 support?

Whisper Large v3 supports transcription in 99 languages and can translate speech from those languages into English.

Is Whisper Large v3 open source?

Yes. Whisper Large v3 is released by OpenAI under the MIT license, meaning the model weights and code are freely available for use and modification.

Does Whisper Large v3 have a context window?

The metadata does not specify a context window for this model. As a speech-to-text model, it processes audio input rather than text tokens, so the standard context window concept does not directly apply.

What is the pricing for Whisper Large v3 on MindStudio?

Specific pricing is not listed in the available metadata, but the model is tagged as low cost. Refer to DeepInfra's pricing page for current per-second or per-request rates.

When was Whisper Large v3 released?

Whisper Large v3 was released in November 2023 by OpenAI.

What types of audio tasks is Whisper Large v3 best suited for?

Whisper Large v3 is designed for automatic speech recognition and translation tasks, including transcribing meetings, generating subtitles, processing multilingual audio content, and building voice interfaces.

Start building with Whisper Large v3

No API keys required. Create AI-powered workflows with Whisper Large v3 in minutes — free.