Wan 3.0
Wan 3.0 is a video generation model that converts text prompts and images into video clips with optional audio output.
Text and image to video with audio
Wan 3.0 is a video generation model published by Wan and served through the WaveSpeed provider on MindStudio. It supports both text-to-video and image-to-video generation, accepting a start image, an end image, reference images, reference videos, and reference audio as inputs, giving creators fine-grained control over the output. The model also includes a thinking mode toggle and native audio generation, making it one of the more input-flexible video generation options available in the catalog.
Wan 3.0 is suited for workflows that require generating short video clips from descriptive prompts or existing visual assets, with configurable aspect ratio, resolution, and duration. Pricing is usage-based at $0.06 to $0.28 per second of generated video, which scales with resolution and duration choices. The model is tagged as the latest release in the Wan series and is available for immediate use on MindStudio without requiring separate API credentials.
What Wan 3.0 supports
Text to Video
Generates video clips directly from text prompts, allowing users to describe a scene and receive a rendered video output.
Image to Video
Animates a provided start image into a video sequence, with optional support for a separate last image to define the ending frame.
Audio Generation
Optionally generates audio alongside the video output via a toggleable setting, supporting reference audio inputs to guide the result.
Reference-Guided Generation
Accepts arrays of reference images and reference videos to steer the visual style or content of the generated clip.
Configurable Output Format
Exposes controls for aspect ratio, resolution, and duration in seconds, letting users tailor the output dimensions to their target format.
Thinking Mode
A toggleable thinking mode that can be enabled to influence how the model interprets and plans the generation from a given prompt.
Reproducible Outputs
Supports a seed input so users can reproduce or iterate on a specific generated video by reusing the same seed value.
Ready to build with Wan 3.0?
Get Started FreeCommon questions about Wan 3.0
How is Wan 3.0 priced on MindStudio?
Wan 3.0 is priced at $0.06 to $0.28 per second of generated video. The exact cost depends on the resolution and duration settings you choose for each generation.
What input types does Wan 3.0 accept?
Wan 3.0 accepts text prompts, a start image, a last image, arrays of reference images, arrays of reference video URLs, and arrays of reference audio URLs. You can also configure aspect ratio, resolution, duration, and a seed value.
Does Wan 3.0 support audio in the generated video?
Yes. Wan 3.0 includes a toggleable audio generation option. You can also supply reference audio files to guide the audio output.
What is the context window for Wan 3.0?
Wan 3.0 has a context window of 50,000 tokens, which governs the length and complexity of the text prompt it can process.
What modes does Wan 3.0 support for video generation?
Wan 3.0 supports at minimum text-to-video and image-to-video modes, selectable via the Mode input. Image-to-video mode uses a start image and optionally a last image to define the clip.
Documentation & links
Parameters & options
First-frame image URL to guide the video generation.
Optional last-frame image URL for video continuation.
Reference image URLs (up to 10). At least one reference image, video, or audio is required.
Reference video URLs (up to 5, total length must not exceed 15 seconds).
Reference audio URLs (up to 5, total length must not exceed 15 seconds).
Duration of the generated video in seconds.
Whether to include audio in the output video.
Enable deep-thinking mode for more deliberate prompt interpretation.
Explore similar models
Start building with Wan 3.0
No API keys required. Create AI-powered workflows with Wan 3.0 in minutes — free.