Veo 3.1
Veo 3.1 is a Google video generation model that creates videos from text prompts or images, with optional audio output.
Text and image to video with audio
Veo 3.1 is a video generation model developed by Google, available under the full model name veo-3.1-generate-preview. It accepts text prompts and images as inputs and produces video outputs, with support for configurable duration, aspect ratio, resolution, and optional audio generation. The model operates with a 2000-token context window and was released in October 2025 as Google's flagship video generation offering.
Veo 3.1 is designed for workflows that require flexible video creation, including text-to-video and image-to-video generation. It supports first-frame and last-frame image conditioning, allowing users to anchor the start or end of a generated clip to a specific image. Additional controls include reference image arrays, negative prompts, seed values for reproducibility, and an auto-enhance prompt option that can refine input descriptions before generation.
What Veo 3.1 supports
Text to Video
Generates video clips directly from text prompts using a 2000-token context window for describing scenes, motion, and style.
Image to Video
Animates a provided image into a video clip, supporting both first-frame and last-frame image conditioning to control clip boundaries.
Audio Generation
Optionally generates synchronized audio alongside the video output, toggled via the Include Audio parameter.
Resolution & Aspect Ratio
Supports configurable resolution and aspect ratio settings, letting users tailor output dimensions to their target format.
Reference Image Input
Accepts an array of reference images to guide visual style or subject consistency across the generated video.
Prompt Auto-Enhancement
An optional auto-enhance mode rewrites or expands the user's text prompt before generation to improve output quality.
Negative Prompting
Accepts a negative prompt to explicitly exclude unwanted visual elements, styles, or content from the generated video.
Reproducible Outputs
Supports a seed parameter so that identical inputs and seed values produce consistent, repeatable video results.
Ready to build with Veo 3.1?
Get Started FreeCommon questions about Veo 3.1
What is the context window for Veo 3.1?
Veo 3.1 has a context window of 2000 tokens, which applies to the text prompt describing the video to be generated.
What input types does Veo 3.1 support?
Veo 3.1 accepts text prompts, single images (for first-frame or last-frame conditioning), and arrays of reference images. It also supports configuration parameters for duration, aspect ratio, resolution, audio, and seed.
Does Veo 3.1 generate audio?
Yes. Audio generation is an optional feature that can be toggled on or off via the Include Audio parameter when making a request.
What is the pricing for Veo 3.1?
Pricing information for Veo 3.1 is not published in the available metadata. You should consult Google Cloud's Vertex AI pricing page for current rates.
When was Veo 3.1 released?
Veo 3.1 was released in October 2025 and is available under the model identifier veo-3.1-generate-preview.
Can I control the length of the generated video?
Yes. Duration is a configurable input parameter, allowing you to select the length of the output video clip.
What people think about Veo 3.1
Community discussion around Veo 3.1 is limited but positive, with a well-received Reddit post in r/midjourney showcasing it alongside Midjourney and Runway in an animated short that earned over 1,300 upvotes. Users in that thread responded favorably to the visual quality of the AI-generated animation.
The thread focused primarily on the creative workflow of combining multiple AI tools rather than on Veo 3.1 in isolation, making it difficult to draw specific conclusions about the model's perceived strengths or limitations from community feedback alone.
Parameters & options
Optional URL of an input image to animate.
Optional URL of the last frame of the video.
Provide up to 10 references images of the scene, subject, objects, or anything else in the image.
Description of what to exclude from an image.
A specific value that is used to guide the 'randomness' of the generation.
Explore similar models
Start building with Veo 3.1
No API keys required. Create AI-powered workflows with Veo 3.1 in minutes — free.