AI Audio: Voice, Speech & Music
AI for audio — real-time voice agents (Pika Me-style), text-to-speech, voice cloning (ElevenLabs), music generation (Suno, Udio), sound effects, audio editing, transcription. Anything where the output or input is audio.

How to Train Your Own TTS Model Locally with Pocket TTS
Kyutai open-sourced the full Pocket TTS training stack. Here's how to train a custom CPU-runnable voice model on your own GPU and data.

Is Breeze TTS 2 Free? License and Commercial Use Explained
Breeze TTS 2's weights are free for research and non-commercial use only. Here's what the license actually allows and how to get commercial rights.

PhoneLLM Cost Per Minute: The Real Economics of Voice Agent LLMs
PhoneLLM's self-hosted cost-per-minute economics on B200 GPUs, benchmarked against API-based voice agent LLMs like GPT 5.6 Terra.

PhoneLLM Alpha 1: Pipecat's Purpose-Built Model for Voice Agents
Pipecat's PhoneLLM Alpha 1, a 30B Nemotron fine-tune for phone voice agents, matches GPT-5.6 Terra accuracy at 94% lower cost and lower latency.

How to Deploy PhoneLLM Alpha 1 with vLLM, SGLang, or Modal
A practical guide to self-hosting PhoneLLM Alpha 1 for voice agents, covering vLLM and SGLang settings, hardware needs, and Modal AutoEndpoints.

Breeze TTS 2: Specs, VRAM Needs, and Local Setup Guide
Breeze TTS 2's specs: sub-40ms latency, 12GB minimum VRAM, voice cloning and design features, and how to run it locally.

Breeze TTS 2: The Open-Weight TTS Model Topping the Leaderboard
Breeze TTS 2 is an open-weight text-to-speech model with sub-40ms latency, voice design, and a #1 spot on the Artificial Analysis leaderboard.

Retell AI Pricing, Free Tier, and How It Builds Voice Agents
How Retell AI's free tier, concurrency limits, and no-code builder work, based on a hands-on build of a phone-based AI voice agent.

How to Run Breeze TTS 2 Locally: GPU Requirements and Setup
Step-by-step guide to self-hosting Breeze TTS 2, covering GPU memory needs, Docker builds, and voice clone, design, and direction commands.

What Is S1 Mini? The Tiny Model That Cleans Up Dictation Text
S1 Mini is a small local model built to strip filler words and fix self-corrections in speech-to-text output. Here's how it works.

What Is HappyShrimp? Alibaba's New AI Music Generator Explained
HappyShrimp is Alibaba's new AI music platform, a fresh challenger to Suno. Here's what it does, how it sounds, and what to know before trying it.

HappyShrimp AI Music Pricing: Free Tier, Paid Plans, and Licensing Gaps
HappyShrimp's free tier, $5 and $20 monthly plans, song limits, and unclear commercial licensing terms, explained for creators considering the tool.

HappyShrimp vs Suno: Which AI Music Generator Sounds Better?
A hands-on look at how Alibaba's new HappyShrimp AI music generator stacks up against Suno v5.5 on vocals, genre range, and anthem rock.

MiniMax Music 3: The Open-Weight AI Music Model, Explained
MiniMax Music 3 is an open-weight AI music generator you can run locally. Here's how it works, what it needs, and how it stacks up to Suno.

MiniMax Music 3 Review: Is This Open AI Music Model Any Good?
A hands-on test of MiniMax Music 3 across pop, Bollywood, and cumbia genres finds solid English vocals but weak multilingual output.

Suno AI's New Download Limits: What Changed and Why It Matters
Suno AI now caps monthly song downloads by plan tier starting September 3rd. Here's what the change means for your music and commercial rights.

Why Voice Tools Like Whispr Flow Signal a New Computing Paradigm
Voice interfaces are moving from novelty to infrastructure. Here's why builders treat Whispr Flow as proof that voice is the next computing layer.

ChatGPT's Voice Mode Overhaul: What Changed and How to Use It
OpenAI rebuilt ChatGPT's interactive voice across web, mobile, and desktop with live video, screen sharing, and work integration.

Claude Voice Mode Arrives: How It Stacks Up Against ChatGPT Voice
Claude's web app finally gets real interactive voice chat. Here's how it compares to ChatGPT's voice overhaul across chat, work, and code.

ElevenLabs Expressive Mode: Giving AI Voice Agents Real Emotion
ElevenLabs' Expressive Mode adds emotional tone control to real-time voice agents, letting builders create character-driven AI personas with memory.

ByteDance Seed 1.0 Audio Gets Precision Dialogue Timestamp Pinning
ByteDance quietly upgraded Seed 1.0 audio generation with timestamp pinning for dialogue, tightening sync for AI voice and sound design work.

GPT Live Voice Mode: Real-Time Translation and Natural Conversation Explained
GPT Live is OpenAI's new full-duplex voice mode that supports real-time translation and natural interruptions. Here's how it works and when to use it.

What Is GPT Live 1? OpenAI's Full-Duplex Voice Model Explained
GPT Live 1 is OpenAI's new conversational voice model with full-duplex interaction, delegation to GPT 5.5, and real-time translation built in.

What Is Seed Audio 1.0? ByteDance's Audio Scene Generator for AI Workflows
Seed Audio 1.0 generates full audio scenes with dialogue, ambient sound, and effects. Learn how it works and how to use it in AI video workflows.