Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Lip Sync Model

Omni Human 1.5

Omni Human 1.5 is a lip sync model from ByteDance that animates a portrait image using an audio input.

PublisherByteDance
TypeLip Sync
Context Window50,000 tokens
ReleasedSeptember 2025
Price$0.25/s
ProviderWaveSpeed
IMAGE+AUDIO

Audio-driven lip sync from a single image

Omni Human 1.5 is a lip sync model developed by ByteDance and made available through the WaveSpeed provider. It takes a single portrait image and an audio file as inputs and generates a video in which the subject's lip movements are synchronized to the provided audio. The model is identified under the full name bytedance/avatar-omni-human-1.5 and was released in September 2025.

The model is designed for use cases that require realistic talking-head video generation without needing a recorded video of the subject speaking. Developers can use it to produce dubbed content, virtual presenters, or animated avatars by supplying only a still image and a corresponding audio track. It is priced at $0.25 per second of output and is available on MindStudio without requiring separate API key management.

What Omni Human 1.5 supports

Lip Sync Generation

Animates a portrait image so that the subject's mouth movements match a provided audio track, producing a talking-head video output.

Image Input

Accepts a single still portrait image via URL as the visual source for the generated animation.

Audio Input

Takes an audio file via URL — such as speech or narration — and uses it to drive the lip and facial animation of the portrait.

Avatar Animation

Generates animated talking-head video from a static image, enabling virtual presenter or dubbed avatar creation without source video footage.

Ready to build with Omni Human 1.5?

Get Started Free

Common questions about Omni Human 1.5

What inputs does Omni Human 1.5 require?

The model requires two inputs: an image URL pointing to a portrait photo and an audio URL pointing to the speech or audio track you want the subject to appear to speak.

How is Omni Human 1.5 priced?

Omni Human 1.5 is priced at $0.25 per second of generated output video.

Does Omni Human 1.5 have a context window?

The model metadata lists a context window of 50,000 tokens, though in practice the primary inputs are an image and an audio file rather than text.

Who developed Omni Human 1.5 and when was it released?

Omni Human 1.5 was developed by ByteDance and released in September 2025. It is served through the WaveSpeed provider on MindStudio.

What is Omni Human 1.5 best suited for?

It is best suited for generating talking-head videos from a single portrait image and an audio file, making it useful for virtual presenters, dubbed avatars, and animated spokesperson content.

Parameters & options

ImageImage URL

Image to be lip synced.

AudioAudio URL

Audio to be lip synced.

Start building with Omni Human 1.5

No API keys required. Create AI-powered workflows with Omni Human 1.5 in minutes — free.