Minimax Speech 2.8 HD
Minimax Speech 2.8 HD is a text-to-speech model from MiniMax supporting emotion, pitch, speed, and multi-format audio output.
High-definition text to speech with emotion control
Minimax Speech 2.8 HD is a text-to-speech model developed by MiniMax and released in January 2026. It accepts up to 50,000 tokens of input text and converts it to audio with configurable voice, speed, pitch, volume, and emotional tone. The model is served through Wavespeed and is priced at $0.007 per run, making it accessible for both prototyping and production use.
The model exposes a range of audio output controls including sample rate, bitrate, channel configuration, and file format selection, giving developers fine-grained control over the resulting audio quality and size. A language boost option and English normalization toggle are also available, which help improve pronunciation accuracy for specific languages and standardize spoken representations of numbers, abbreviations, and symbols. These controls make it well-suited for applications such as voiceovers, audiobook generation, accessibility tools, and multilingual content pipelines.
What Minimax Speech 2.8 HD supports
Emotion Control
Allows selection of an emotional tone for synthesized speech, enabling outputs that sound expressive rather than flat or robotic.
Voice Customization
Supports adjustable speed, pitch, and volume parameters alongside a selectable voice ID to tailor the output voice to specific use cases.
Audio Format Options
Outputs audio in selectable formats with configurable sample rate, bitrate, and mono or stereo channel settings.
Language Boost
Provides a language boost input that improves pronunciation accuracy for targeted languages beyond the default behavior.
English Normalization
A toggle that standardizes how numbers, abbreviations, and symbols are spoken aloud in English-language text.
Large Context Input
Accepts up to 50,000 tokens of input text, supporting long-form content such as articles, scripts, or book chapters in a single request.
Ready to build with Minimax Speech 2.8 HD?
Get Started FreeCommon questions about Minimax Speech 2.8 HD
How much does it cost to use Minimax Speech 2.8 HD?
The model is priced at $0.007 per run, regardless of the length of the input text up to the context limit.
What is the maximum input length for this model?
Minimax Speech 2.8 HD supports a context window of 50,000 tokens, which is sufficient for long documents, scripts, or articles.
What audio output formats and quality settings are available?
The model lets you select the output file format, sample rate, bitrate, and channel configuration (mono or stereo), giving you control over audio quality and file size.
Can I control how the voice sounds beyond just choosing a voice ID?
Yes. You can adjust speed, pitch, and volume numerically, and you can also select an emotion to influence the expressive quality of the synthesized speech.
Does the model support languages other than English?
The model includes a language boost input designed to improve pronunciation accuracy for specific languages, and an English normalization toggle for handling numbers and abbreviations in English text.
Documentation & links
Parameters & options
Voice preset to use for speech synthesis.
Speech speed multiplier.
Volume level.
Pitch adjustment.
Emotional tone of the speech delivery.
Audio sample rate in Hz.
Audio bitrate in bits per second.
Audio channel configuration.
Output audio format.
Boost recognition for a specific language.
Improves number-reading performance in English text (dates, currencies, etc.).
Explore similar models
Start building with Minimax Speech 2.8 HD
No API keys required. Create AI-powered workflows with Minimax Speech 2.8 HD in minutes — free.