Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Speech to Text Model

Voxtral Mini 3B

Voxtral Mini 3B is a 3-billion-parameter open-source speech-to-text model released by Mistral in July 2025.

PublisherMistral
TypeTranscription
ReleasedJuly 2025
Input$16.67/MTok
OutputFree/MTok
ProviderDeepInfra
OPEN SOURCE

Compact open-source audio transcription from Mistral

Voxtral Mini 3B is a speech-to-text model developed by Mistral and released in July 2025 under the model identifier mistralai/Voxtral-Mini-3B-2507. It is a 3-billion-parameter model designed for audio transcription tasks and is made available as an open-source release, meaning the weights are publicly accessible for research and deployment. The model is served through DeepInfra, making it accessible via API without requiring users to manage their own infrastructure.

Voxtral Mini 3B is suited for developers and teams that need a compact, open-source transcription model they can run or fine-tune without the overhead of larger architectures. Its smaller parameter count makes it a practical option for latency-sensitive or resource-constrained environments where a full-scale model would be impractical. Being open source, it can also be adapted and deployed on private infrastructure for use cases where data privacy or customization is a priority.

What Voxtral Mini 3B supports

Audio Transcription

Converts spoken audio into text. Designed specifically for transcription tasks as a dedicated speech-to-text model.

Open Source Weights

Model weights are publicly available under Mistral's open-source release, allowing self-hosting, fine-tuning, and private deployment.

Compact 3B Architecture

Built on a 3-billion-parameter architecture, keeping resource requirements lower than larger transcription models while retaining usable accuracy.

API Access via DeepInfra

Served through DeepInfra's inference infrastructure, enabling API-based transcription without requiring users to host the model themselves.

Fine-Tuning Ready

As an open-source model, Voxtral Mini 3B can be fine-tuned on domain-specific audio datasets to improve transcription accuracy for specialized vocabularies.

Ready to build with Voxtral Mini 3B?

Get Started Free

Common questions about Voxtral Mini 3B

What type of model is Voxtral Mini 3B?

Voxtral Mini 3B is a speech-to-text (transcription) model. It takes audio input and produces text output.

Who made Voxtral Mini 3B and when was it released?

Voxtral Mini 3B was developed by Mistral and released in July 2025 under the identifier mistralai/Voxtral-Mini-3B-2507.

Is Voxtral Mini 3B open source?

Yes. Voxtral Mini 3B is tagged as open source, meaning the model weights are publicly available for download, fine-tuning, and self-hosted deployment.

What is the context window for Voxtral Mini 3B?

The context window for Voxtral Mini 3B is not specified in the available metadata. Check the DeepInfra model page or Mistral's documentation for the latest details.

How is Voxtral Mini 3B priced?

Pricing information is not included in the available metadata. Refer to the DeepInfra model page at deepinfra.com/mistralai/Voxtral-Mini-3B-2507 for current pricing.

What is the knowledge cutoff for Voxtral Mini 3B?

Voxtral Mini 3B is a transcription model rather than a language model, so a knowledge cutoff date is not applicable in the traditional sense. The model was released in July 2025.

Start building with Voxtral Mini 3B

No API keys required. Create AI-powered workflows with Voxtral Mini 3B in minutes — free.