Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Lip Sync Model

AI Avatar Standard

AI Avatar Standard is a lip sync model from Kling that animates a portrait image to match a provided audio track.

PublisherKling
TypeLip Sync
Context Window50,000 tokens
ReleasedAugust 2025
Price$0.28+/s (dynamic)
ProviderWaveSpeed
IMAGE+AUDIO

Lip sync video generation from image and audio

AI Avatar Standard is a lip sync model developed by Kling and served through the Wavespeed provider on MindStudio. It takes a portrait image and an audio file as inputs and generates a video in which the subject's mouth movements are synchronized to the spoken audio. Users can also supply a text prompt and select an output resolution, with a seed parameter available for reproducible results.

The model is suited for use cases such as creating talking avatar videos, dubbing static images with voiceover, and producing animated spokesperson content without requiring video footage of the original subject. Pricing is dynamic and starts at $0.28 per second of generated video, making cost proportional to the length of the output clip. It was released in August 2025 and is currently available on MindStudio under the model ID kwaivgi/kling-v1-ai-avatar-standard.

What AI Avatar Standard supports

Lip Sync Animation

Animates a portrait image so that mouth movements match a supplied audio track, producing a synchronized talking-head video.

Image Input

Accepts a portrait image URL as the visual base for the generated avatar video.

Audio Input

Takes an audio URL as the speech source that drives lip sync timing and mouth shape generation.

Resolution Selection

Allows users to choose the output video resolution via a select input, giving control over output quality and file size.

Seed Control

Supports a seed parameter so that generation results can be reproduced consistently across multiple runs.

Prompt Guidance

Accepts a text prompt to provide additional context or stylistic direction alongside the image and audio inputs.

Ready to build with AI Avatar Standard?

Get Started Free

Common questions about AI Avatar Standard

What inputs does AI Avatar Standard require?

The model requires an image URL (portrait) and an audio URL. A text prompt, resolution selection, and seed value are optional additional inputs.

How is AI Avatar Standard priced?

Pricing is dynamic and starts at $0.28 per second of generated video output, so the total cost scales with the length of the audio clip provided.

Does AI Avatar Standard have a context window?

The model metadata lists a context window of 50,000 tokens, though for a lip sync model this primarily governs any text-based prompt or metadata processing rather than video length directly.

What resolution options are available?

Resolution is configurable via a select input at generation time. Specific resolution values should be confirmed in the Wavespeed or MindStudio API documentation for this model.

When was AI Avatar Standard released?

AI Avatar Standard was released in August 2025 and is currently available on MindStudio under the full model name kwaivgi/kling-v1-ai-avatar-standard.

Parameters & options

ImageImage URL

Image to be lip synced.

AudioAudio URL

Audio to be lip synced.

PromptPrompt

Optional prompt to guide the lip sync.

ResolutionSelect

The resolution of the output video.

Default: 480p
480p (default)720p
SeedSeed
Default: -1

Start building with AI Avatar Standard

No API keys required. Create AI-powered workflows with AI Avatar Standard in minutes — free.