A Start-to-End Frame Video Model for High-Quality Shot Creation
veo31 is the quality-first video model in the Google Veo 3.1 series, suitable for turning text concepts, product stills, and storyboards into dynamic shots. It supports text-to-video, image-to-video with a first frame or start-and-end frames, and can further extend generated clips. Compared with Fast models designed primarily for rapid drafting, veo31 is better suited for asset production and shot review where image quality takes priority.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
1 first frame; 2 start-and-end frames, submitted via image_urls
Output resolution
Default 720p; resolution options include 1080p, 4k, gif
Aspect ratios
16:9 landscape, 9:16 portrait
Video extension
/veo/extend; 720p or 1080p
Prompt processing
Text descriptions; translation can enable automatic translation
Result delivery
Video link and video ID; supports asynchronous tasks, callbacks, and get1080p
The above are this platform's veo31 invocation specifications; generation and extension each use their own resolution options, and series capabilities should not be regarded as jointly supported across all operations.
Core Capabilities
Define storyboard start and end points more clearly
Image-to-video can begin with one first-frame image, or use two images to specify the start and end points. Compared with text alone, this approach establishes the composition at both ends of the shot first, then uses prompts to describe the action in between, camera movement, and the transition process. It is suitable for product reveals, scene transitions, and storyboard previsualization.
Quality-first visual production
veo31 belongs to Quality mode and is suitable for tasks where review focuses on visual presentation rather than rapid drafting. During generation, you can select landscape or portrait aspect ratios and resolution options such as 1080p and 4k, allowing the same concept to serve both landscape displays and portrait content; compositions should be redesigned for the target aspect ratio.
Continue creating from generated shots
Generated clips can be further developed through extension, with new prompts guiding subsequent visuals—for example, having the camera pull back or the subject continue moving. Asynchronous tasks and callbacks make it easy to connect generation, review, extension, and asset archiving into a workflow, without requiring the application to wait continuously for video generation requests to finish.
Use Cases
Product Showcases and Advertising Shots
Use a product still as the first frame, describe subject movement, camera angles, and lighting changes, and create dynamic assets for advertising edits. If you already have a clear ending composition, add last-frame constraints to define the shot endpoint. After delivery, focus on checking the product shape, branding, and details, then combine it with subtitles and music.
Storyboard Shots and Scene Transitions
Submit the opening and ending images from a storyboard as the first and last frames, along with descriptions of character actions, spatial changes, and camera direction, to generate transition shots for discussion. It is suitable for verifying whether visual continuity works, helping directors and design teams evaluate shot options before formal production, rather than directly replacing production of an entire film.
Shot Continuation and Asset Expansion
Select generated clips that have passed review, use their video IDs to initiate extensions, and describe the actions or shot scale desired in the next segment. The returned video links can enter the editing asset library; extension results can also be extended further. Keep the original clips during production to facilitate continuity comparison and fallback options.
How to Choose This Model
Choose veo31 for Quality First, Consider Fast for Drafts
When a task already has a clear composition and you want to conduct quality review around a small number of options, prioritize veo31. If you need to first explore a large number of prompts and camera directions, consider veo31-fast, then select options for quality-first production. The two are independent models; when comparing them, keep the image, aspect ratio, and prompt consistent, and do not assume a fixed speed difference.
Choose First/Last Frames and Multi-Image Fusion Separately
If the task is to animate an existing image or create a transition between specified starting and ending images, veo31's first-frame and first/last-frame modes are more direct. If you want to blend elements from multiple assets into a single creation, consider veo31-fast-ingredients. The latter uses multi-image fusion and requires image uploads; it is not equivalent to veo31's first/last-frame control.
Getting Started
Choose Text or Starting/Ending Images
Use text2video for text only; use image2video for image-driven generation. One image in image_urls is the first frame, and two images are the first and last frames; describe the action that occurs in between.
Choose the Dedicated Creation Endpoint
Specify model=veo31 for /veo/videos, choose the action based on text or images, and fill in prompt. Start with aspect_ratio=16:9 and resolution=720p, then increase the output tier within this model's supported range.
Save Task IDs and Video IDs Separately
After setting async=true, save the task_id and obtain the finished video through /veo/tasks or a callback; also save data[].id. Veo 3.1 extensions use video IDs and cannot use task IDs as a substitute.
Trial Recommendations: First/Last Frames and Extension
Input and Goal
The first frame is a closed wooden door, and the last frame is an open wooden door with a warm interior. The shot slowly moves in from outside the door, keeping the doorframe and environment consistent.
Acceptance and Next Steps
First verify the perspective and composition of the two images, then check the motion in between; for subsequent extension, save data[].id. Generating in 4k does not mean extension also supports 4k.
Usage Limits
veo31 image-to-video is designed for up to two images: one for the first frame, or two for the first and last frames. Do not use multiple assets directly as fusion references; when the start and end images differ significantly, simplify the intermediate motion and check subject shape and spatial continuity.
Generation and extension have different resolution ranges: generation can be 4k, while extension only provides 720p and 1080p. Extension should use data[].id from the current generation or extension result, rather than task_id; video IDs from earlier generations may not be usable for extension.
Extension results can be extended further, but cannot be sent to /veo/reshoot to adjust camera motion or /veo/objects to add or remove objects. If these editing operations are needed later, retain the original video and plan the operation order in advance to avoid saving only the final extended video.
Frequently Asked Questions
Can veo31 generate videos using only text?
Yes. Use /veo/videos, set model to veo31, action to text2video, and provide a scene prompt. It is recommended to clearly specify the subject, action, environment, and camera direction; if the composition is already determined, you can instead use image2video, using an image as the starting point of the shot.
How should one image and two images be used respectively?
One image is used as the first frame, so the video starts from the specified image; two images are used as the first and last frames, guiding the shot to connect the start and end points. Submit image links through image_urls, then use prompt to describe the transition. veo31 should not be used as a three-image blending method.
Can the 4k option in veo31 also be used for extensions?
When generating a video, you can set resolution to 4k, but the resolution options for /veo/extend are 720p or 1080p; the two operations cannot share the same set of settings. If the project requires a consistent delivery resolution, determine the extension requirements first, then arrange asset generation and final output.
How do I continue an already generated video?
Submit model=veo31 and data[].id from the generation result to /veo/extend, and you can add a prompt to guide the subsequent footage. This uses the video ID, not the asynchronous task's task_id. The new video ID obtained after the extension is complete can also be used for the next extension.
How do Chinese prompts work with asynchronous generation?
You can describe scenes in Chinese and enable automatic translation through translation. After setting async=true, first obtain the task_id, then query the task result; you can also set callback_url to receive completion notifications. The video_url in the completed result is used to retrieve the video, and data[].id is used for subsequent video operations.