LatentSync
LatentSync is a lip sync model from ByteDance that synchronizes video facial movements to a provided audio track.
Audio-driven lip sync for video content
LatentSync is a lip sync model developed by ByteDance and made available through the Wavespeed provider on MindStudio. It takes an audio file and a video file as inputs and generates output video in which the subject's lip movements are synchronized to the provided audio. The model is categorized under audio and video processing and is listed as available starting at $0.15 per use.
LatentSync is designed for workflows that require automated dubbing, voice-over alignment, or any scenario where existing video footage needs to match a different or modified audio track. Because it accepts standard audio and video URLs as inputs, it can be integrated into pipelines that handle media transformation tasks. It is suited for content localization, post-production editing, and AI-generated video applications where lip accuracy to audio is a requirement.
What LatentSync supports
Lip Sync Generation
Synchronizes the mouth and lip movements of a subject in a video to match a provided audio track. Accepts audioUrl and videoUrl inputs to produce aligned output video.
Audio Input Processing
Accepts an external audio file via URL as the reference track for driving lip movement generation. Supports audio-driven animation without requiring manual keyframing.
Video Input Processing
Takes an existing video file via URL as the source footage for lip sync transformation. The original video frames are used as the base for facial region modification.
Media Pipeline Integration
Designed to accept URL-based inputs for both audio and video, making it compatible with automated media processing workflows. Can be embedded in dubbing, localization, or post-production pipelines.
Ready to build with LatentSync?
Get Started FreeCommon questions about LatentSync
What inputs does LatentSync require?
LatentSync requires two inputs: an audio file URL and a video file URL. The model uses the audio track to drive lip movement synchronization on the subject in the provided video.
What does LatentSync cost to use?
LatentSync is priced starting at $0.15 per use on MindStudio. Actual costs may vary depending on the length or complexity of the media being processed.
Who developed LatentSync?
LatentSync was developed by ByteDance and is served through the Wavespeed provider on MindStudio.
What is LatentSync best used for?
LatentSync is best suited for dubbing, voice-over alignment, content localization, and post-production workflows where video footage needs to be synchronized to a different or modified audio track.
Does LatentSync have a context window or knowledge cutoff?
LatentSync is a video and audio processing model, not a language model, so a knowledge cutoff date does not apply. It has a listed context window value of 50,000 tokens in the catalog metadata, though this parameter is more relevant to text-based models.
Parameters & options
Audio to be synchronized.
Video to be synchronized.
Explore similar models
Start building with LatentSync
No API keys required. Create AI-powered workflows with LatentSync in minutes — free.