Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Text to Speech Model

Minimax Speech 2.8 HD

Minimax Speech 2.8 HD is a text-to-speech model from MiniMax supporting emotion, pitch, speed, and multi-format audio output.

PublisherMiniMax
TypeText to Speech
Context Window50,000 tokens
ReleasedJanuary 2026
Price$0.007/run
ProviderWaveSpeed
SPEECH

High-definition text to speech with emotion control

Minimax Speech 2.8 HD is a text-to-speech model developed by MiniMax and released in January 2026. It accepts up to 50,000 tokens of input text and converts it to audio with configurable voice, speed, pitch, volume, and emotional tone. The model is served through Wavespeed and is priced at $0.007 per run, making it accessible for both prototyping and production use.

The model exposes a range of audio output controls including sample rate, bitrate, channel configuration, and file format selection, giving developers fine-grained control over the resulting audio quality and size. A language boost option and English normalization toggle are also available, which help improve pronunciation accuracy for specific languages and standardize spoken representations of numbers, abbreviations, and symbols. These controls make it well-suited for applications such as voiceovers, audiobook generation, accessibility tools, and multilingual content pipelines.

What Minimax Speech 2.8 HD supports

Emotion Control

Allows selection of an emotional tone for synthesized speech, enabling outputs that sound expressive rather than flat or robotic.

Voice Customization

Supports adjustable speed, pitch, and volume parameters alongside a selectable voice ID to tailor the output voice to specific use cases.

Audio Format Options

Outputs audio in selectable formats with configurable sample rate, bitrate, and mono or stereo channel settings.

Language Boost

Provides a language boost input that improves pronunciation accuracy for targeted languages beyond the default behavior.

English Normalization

A toggle that standardizes how numbers, abbreviations, and symbols are spoken aloud in English-language text.

Large Context Input

Accepts up to 50,000 tokens of input text, supporting long-form content such as articles, scripts, or book chapters in a single request.

Ready to build with Minimax Speech 2.8 HD?

Get Started Free

Common questions about Minimax Speech 2.8 HD

How much does it cost to use Minimax Speech 2.8 HD?

The model is priced at $0.007 per run, regardless of the length of the input text up to the context limit.

What is the maximum input length for this model?

Minimax Speech 2.8 HD supports a context window of 50,000 tokens, which is sufficient for long documents, scripts, or articles.

What audio output formats and quality settings are available?

The model lets you select the output file format, sample rate, bitrate, and channel configuration (mono or stereo), giving you control over audio quality and file size.

Can I control how the voice sounds beyond just choosing a voice ID?

Yes. You can adjust speed, pitch, and volume numerically, and you can also select an emotion to influence the expressive quality of the synthesized speech.

Does the model support languages other than English?

The model includes a language boost input designed to improve pronunciation accuracy for specific languages, and an English normalization toggle for handling numbers and abbreviations in English text.

Parameters & options

VoiceSelect

Voice preset to use for speech synthesis.

Default: Friendly_Person
Wise WomanFriendly PersonInspirational GirlDeep Voice ManCalm WomanCasual GuyLively GirlPatient ManYoung KnightDetermined ManLovely GirlDecent BoyImposing MannerElegant ManAbbessSweet Girl 2Exuberant Girl
SpeedNumber

Speech speed multiplier.

Default: 1Range: 0.5–2 (step 0.1)
VolumeNumber

Volume level.

Default: 1Range: 0–2 (step 0.1)
PitchNumber

Pitch adjustment.

Default: 0Range: -12–12 (step 1)
EmotionSelect

Emotional tone of the speech delivery.

HappySadAngryFearfulDisgustedSurprisedNeutral
Sample RateSelect

Audio sample rate in Hz.

Default: 44100
16,000 Hz24,000 Hz32,000 Hz44,100 Hz (default)
BitrateSelect

Audio bitrate in bits per second.

Default: 128000
32,00064,000128,000 (default)256,000
ChannelSelect

Audio channel configuration.

MonoStereo
FormatSelect

Output audio format.

MP3WAVFLACOGGPCM
Language BoostSelect

Boost recognition for a specific language.

AutoAfrikaansArabicBulgarianCatalanChineseChinese (Yue)CroatianCzechDanishDutchEnglishFilipinoFinnishFrenchGermanGreekHebrewHindiHungarianIndonesianItalianJapaneseKoreanMalayNorwegianNynorskPersianPolishPortugueseRomanianRussianSlovakSlovenianSpanishSwedishTamilThaiTurkishUkrainianVietnamese
English NormalizationToggle Group

Improves number-reading performance in English text (dates, currencies, etc.).

Default: false
DisabledEnabled

Start building with Minimax Speech 2.8 HD

No API keys required. Create AI-powered workflows with Minimax Speech 2.8 HD in minutes — free.