Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Lip Sync Model

LatentSync

LatentSync is a lip sync model from ByteDance that synchronizes video facial movements to a provided audio track.

PublisherByteDance
TypeLip Sync
Context Window50,000 tokens
ReleasedAugust 2025
Price$0.15+
ProviderWaveSpeed
AUDIO+VIDEO

Audio-driven lip sync for video content

LatentSync is a lip sync model developed by ByteDance and made available through the Wavespeed provider on MindStudio. It takes an audio file and a video file as inputs and generates output video in which the subject's lip movements are synchronized to the provided audio. The model is categorized under audio and video processing and is listed as available starting at $0.15 per use.

LatentSync is designed for workflows that require automated dubbing, voice-over alignment, or any scenario where existing video footage needs to match a different or modified audio track. Because it accepts standard audio and video URLs as inputs, it can be integrated into pipelines that handle media transformation tasks. It is suited for content localization, post-production editing, and AI-generated video applications where lip accuracy to audio is a requirement.

What LatentSync supports

Lip Sync Generation

Synchronizes the mouth and lip movements of a subject in a video to match a provided audio track. Accepts audioUrl and videoUrl inputs to produce aligned output video.

Audio Input Processing

Accepts an external audio file via URL as the reference track for driving lip movement generation. Supports audio-driven animation without requiring manual keyframing.

Video Input Processing

Takes an existing video file via URL as the source footage for lip sync transformation. The original video frames are used as the base for facial region modification.

Media Pipeline Integration

Designed to accept URL-based inputs for both audio and video, making it compatible with automated media processing workflows. Can be embedded in dubbing, localization, or post-production pipelines.

Ready to build with LatentSync?

Get Started Free

Common questions about LatentSync

What inputs does LatentSync require?

LatentSync requires two inputs: an audio file URL and a video file URL. The model uses the audio track to drive lip movement synchronization on the subject in the provided video.

What does LatentSync cost to use?

LatentSync is priced starting at $0.15 per use on MindStudio. Actual costs may vary depending on the length or complexity of the media being processed.

Who developed LatentSync?

LatentSync was developed by ByteDance and is served through the Wavespeed provider on MindStudio.

What is LatentSync best used for?

LatentSync is best suited for dubbing, voice-over alignment, content localization, and post-production workflows where video footage needs to be synchronized to a different or modified audio track.

Does LatentSync have a context window or knowledge cutoff?

LatentSync is a video and audio processing model, not a language model, so a knowledge cutoff date does not apply. It has a listed context window value of 50,000 tokens in the catalog metadata, though this parameter is more relevant to text-based models.

Parameters & options

AudioAudio URL

Audio to be synchronized.

VideoVideo URL

Video to be synchronized.

Start building with LatentSync

No API keys required. Create AI-powered workflows with LatentSync in minutes — free.