Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Text to Speech Model

Qwen3 TTS

Qwen3 TTS is an open-source multilingual text-to-speech model from Qwen with a 10,000-token context window.

PublisherQwen
TypeText to Speech
Context Window10,000 tokens
Released2026
Input$20.00/MTok
OutputFree/MTok
ProviderDeepInfra
OPEN SOURCEMULTILINGUAL

Open-source multilingual text to speech

Qwen3 TTS is a text-to-speech model developed by Qwen and served via DeepInfra. It converts text input into spoken audio and supports multiple languages, making it suitable for applications that require voice output across different linguistic contexts. The model is open source and carries a 10,000-token context window, which allows it to handle longer passages of text in a single request.

Qwen3 TTS is designed for developers building voice interfaces, narration tools, accessibility features, or any application that requires synthesized speech. Its multilingual support means it can be used in products targeting audiences who speak different languages without switching to a separate model. Because it is open source, teams can inspect, fine-tune, or self-host the model depending on their deployment requirements.

What Qwen3 TTS supports

Text to Speech

Converts written text into synthesized audio output. Accepts up to 10,000 tokens of input per request, enabling longer passages to be processed in a single call.

Multilingual Support

Generates speech in multiple languages from a single model. This removes the need to maintain separate TTS models for different language targets.

Long Context Input

Supports a 10,000-token context window, allowing full articles, scripts, or documents to be synthesized without chunking.

Open Source

The model weights are publicly available, enabling self-hosting, inspection, and fine-tuning by developers and researchers.

Ready to build with Qwen3 TTS?

Get Started Free

Common questions about Qwen3 TTS

What is the context window for Qwen3 TTS?

Qwen3 TTS supports a context window of 10,000 tokens, which means you can submit longer text passages in a single request without needing to split them manually.

What languages does Qwen3 TTS support?

Qwen3 TTS is tagged as multilingual, meaning it can generate speech in multiple languages. For the exact list of supported languages, refer to the official Qwen model documentation or the DeepInfra model page.

Is Qwen3 TTS open source?

Yes. Qwen3 TTS is tagged as open source, meaning the model weights are publicly available. Developers can inspect, fine-tune, or self-host the model.

What is the pricing for Qwen3 TTS on MindStudio?

Pricing information is not published in the current metadata. Check the MindStudio platform or the DeepInfra provider page for up-to-date pricing details.

What is the knowledge cutoff or release date for Qwen3 TTS?

The metadata lists a release date of 2026. No specific training data cutoff date is available in the current metadata.

What output format does Qwen3 TTS produce?

Qwen3 TTS is a text-to-speech model that produces audio output from text input. Specific audio format details such as sample rate or encoding are not specified in the current metadata; consult the DeepInfra model page for technical output specifications.

Start building with Qwen3 TTS

No API keys required. Create AI-powered workflows with Qwen3 TTS in minutes — free.