Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Video Generation Model

Wan 2.7

Wan 2.7 is a video generation model from Alibaba that supports text-to-video, image-to-video, video editing, and audio generation.

PublisherWan
TypeVideo
Context Window2,000 tokens
ReleasedFebruary 2026
Price$0.05-$0.15/second
ProviderAlibaba Cloud
Source ImageReference ImagesVideo EditingAudio

Text and image-to-video with audio generation

Wan 2.7 is a video generation model published by Wan and provided through Alibaba, released in February 2026. It accepts text prompts, source images, source videos, and reference images as inputs, and can produce video output with optional generated audio. The model supports multiple operational modes, selectable resolutions, aspect ratios, and durations, giving users direct control over the output format.

Wan 2.7 is designed for workflows that require flexible video creation, including generating video from a first-frame image, editing existing video using reference images, and adding audio to generated clips. Pricing is usage-based at $0.05–$0.15 per second of generated video, and the model includes an auto-enhance prompt option that can expand or refine input prompts before generation. It is suited for content creators, developers, and teams building video generation pipelines on MindStudio.

What Wan 2.7 supports

Text-to-Video

Generates video clips from text prompts, with an optional auto-enhance prompt toggle that expands the input before generation.

Image-to-Video

Uses a provided first-frame image as the starting point for video generation, anchoring the visual content to a specific source image.

Video Editing

Accepts a source video and edit reference images to modify or restyle existing footage within the generation pipeline.

Reference Image Support

Accepts one or more reference images to guide visual style or subject consistency across the generated video.

Audio Generation

Includes a toggleable option to generate audio alongside the video output, producing synchronized sound without a separate step.

Resolution & Aspect Ratio Control

Lets users select output resolution and aspect ratio via toggle groups, supporting different formats for various delivery targets.

Shot Type Selection

Provides a shot type toggle that lets users specify the camera framing style, such as close-up or wide shot, for the generated video.

Ready to build with Wan 2.7?

Get Started Free

Common questions about Wan 2.7

How is Wan 2.7 priced on MindStudio?

Wan 2.7 is priced at $0.05–$0.15 per second of generated video. The exact rate within that range depends on the selected mode, resolution, and duration.

What input types does Wan 2.7 accept?

The model accepts text prompts, a first-frame image URL, a source video URL, and one or more reference image arrays. You can also configure mode, resolution, aspect ratio, duration, shot type, and audio generation as part of each request.

Does Wan 2.7 have a context window limit?

Yes, the model has a context window of 2,000 tokens, which applies to the text prompt and associated input parameters.

Can Wan 2.7 generate audio along with video?

Yes. The model includes a toggleable audio generation option that produces audio synchronized with the video output in a single generation step.

Who publishes Wan 2.7 and when was it released?

Wan 2.7 is published by Wan and provided through Alibaba. It was released in February 2026 and became available on MindStudio in June 2026.

Parameters & options

ModeSelect
Default: text-to-video
Text to VideoImage to VideoReference to VideoVideo Edit
First Frame ImageImage URLimage-to-video only

Image used as the first frame of the generated video.

Source VideoVideo URLvideo-edit only

The video to edit. Describe your edits in the prompt using natural language instructions. Supports both localized and global edits.

Reference ImagesImage URL Arrayreference-to-video only

Provide reference images of people or objects to keep their appearance (and voice) consistent in the generated video. Reference them as "Character 1", "Character 2", etc. in the prompt. Supports multi-character joint performances.

Reference ImagesImage URL Arrayvideo-edit only

Optionally provide reference images to seamlessly replace elements in the video.

ResolutionToggle Group
Default: 720P
720p1080p
Aspect RatioToggle Grouptext-to-video, reference-to-video only
Default: 16:9
16:99:16
DurationSelecttext-to-video, image-to-video, reference-to-video only
Default: 5
5s10s15s
Generate AudioToggle Grouptext-to-video, image-to-video, reference-to-video only

Generate synchronized audio including dialogue, vocals, and sound effects alongside the video.

Default: false
YesNo
Shot TypeToggle Grouptext-to-video, image-to-video, reference-to-video only

Multi-shot uses intelligent shot scheduling to generate multi-shot narrative videos with consistent subjects, scenes, and atmosphere.

Default: single
Single ShotMulti Shot
Auto-Enhance PromptToggle Grouptext-to-video, image-to-video, video-edit only

Automatically rewrite and expand the prompt for improved output quality.

Default: true
YesNo

Start building with Wan 2.7

No API keys required. Create AI-powered workflows with Wan 2.7 in minutes — free.