How should I choose between veo3 and veo3-fast?
veo3 is the Quality mode, suitable for creating assets where shot performance matters; veo3-fast is geared more toward rapid iteration, making it suitable for exploring ideas and testing prompts. When choosing, first consider the task stage: prioritize Fast for concept experiments, and consider the standard version for final shots. Do not assume there is always the same difference in processing time.
What is the difference between one image and two images?
Set action to image2video and submit image links through image_urls. One image is used for first-frame creation, while two images are used for first-and-last-frame creation. The prompt should describe the action and camera changes from the starting point to the ending point, rather than treating the two images as a collection of assets that can be freely blended.
Can it generate dialogue and ambient sound?
Veo 3 has native joint audio-video generation capabilities and can create character dialogue, lip synchronization, and scene sound effects. It is recommended to clearly specify the speaking characters, lines, and background sounds in the prompt, then listen and check the result. Automatic prompt translation is a text-processing feature and is not equivalent to video dialogue dubbing or language conversion.
How can I get 1080p video? Does it support 4K?
You can select resolution=1080p, or use action=get1080p for an already generated video. The latter method requires submitting the video ID from the result as video_id, not task_id. veo3 does not support 4K; if you have higher requirements for output clarity, consider veo31.
How can I wait for veo3 generation to complete in an application?
Set async=true to first obtain task_id, then query the task result; you can also set callback_url to receive a completion notification. Successful results include a video link, video ID, and status. Applications should display the video based on the completion status and distinguish between the task ID and video ID, which are used for tracking tasks and obtaining the high-definition version respectively.