All models

wan3.0-video

AlibabaVideo
Get your API key
wan3.0-video

Define complete video shots with text and multimedia assets

wan3.0-video is a video generation model in the Alibaba Wan 3 series, suitable for creating shots from text concepts, first and last frame designs, or multimedia references. It brings different creative approaches together in unified asset inputs, generating videos with specified or smart durations while allowing selection of resolution, aspect ratio, and sound. Through this platform, shot drafts, product showcases, and reference-driven creations can be integrated into asynchronous workflows.

AlibabaModel brand
VideoModel type
Text · Image guidanceCreation method
STANDARD APIs · QUICK SETUP

Bring this model into your workflow

Submit requests to the public API at api.acedata.cloud using the documented parameters, then use the results in your application.

API host
api.acedata.cloud
model
wan3.0-video
Get your API key

Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.

Specifications and API features

Generation methods
Text generation, first and last frame generation, multimedia reference generation
Platform duration
2–30 second integers; duration=-1 for smart duration
Platform resolution
480P、720P、1080P
Output aspect ratio
adaptive、16:9、4:3、1:1、3:4、9:16
Asset inputs
media supports seven asset types, with a maximum total of 10 items
Sound control
audio controls generated sound, default false; reference audio assets supported
Delivery method
Asynchronous task; returns video link, dimensions, and thumbnail information

The figures and controls above are the wan3.0-video invocation specifications for this platform; the Prime accelerated version is a different model.

Core capabilities

From descriptions to asset-driven creation

In addition to describing scenes with prompts, you can submit reference images, videos, and audio to translate appearance, actions, and pacing requirements into specific assets. Wan 3 identifies the creation mode based on media, making it suitable for projects with an existing visual direction, while text can focus on explaining which elements should be retained and which content should be regenerated.

Define shot boundaries with first and last frames

When the opening and closing images are already determined, you can provide first_frame and last_frame separately, then describe the motion and camera changes in between. This approach is suitable for product reveals, subject movement, and transition drafts, keeping creation centered on clearly defined start and end images; it is a different mode from reference asset generation and requests must be organized separately.

Control video according to delivery goals

Duration, resolution, aspect ratio, and sound can be selected in combination based on publishing goals. First validate shots at 480P, then create delivery versions at 720P or 1080P; landscape, portrait, and square content can each use their own aspect ratio. When sound is needed, explicitly enable audio; reference audio should be specified separately as an asset.

Use Cases

Product Showcases and Advertising Shots

Provide product images and camera movement references, and specify the subject, environment, presentation sequence, and visual characteristics to preserve in the prompt to generate product reveal or demonstration clips. If the opening and closing compositions are already designed, use first-and-last-frame mode instead to deliver video assets ready for editing and review.

Social Content and Concept Previsualization

Start with a scene description, clearly define character actions, environmental atmosphere, and camera position, then choose an aspect ratio and duration suited to the publishing channel. Create low-resolution drafts during the concept stage, then adjust output specifications after confirming the direction. This is suitable for turning copywriting ideas into watchable shot plans rather than directly producing a complete long-form video.

Reference-Driven Creative Pipeline

Submit image, video, or audio assets together with creative instructions for style exploration, action references, and pacing previsualization. The application saves the asynchronous task_id, queries results through /wan/tasks, then archives successful videos and thumbnails to the asset library, connecting manual review, editing, and publishing workflows.

How to Choose This Model

Trade-offs with Wan 2.6 Calling Methods

When a new project needs to switch between text, first-and-last-frame, and mixed references, wan3.0-video's unified media input makes assets easier to organize. Existing Wan 2.6 integrations can continue using their original task models and parameters such as action. During migration, redesign the asset structure rather than simply replacing the model name, and do not treat old parameters as the creative capabilities of the new model.

How to Choose Between Standard and Prime

wan3.0-video can be used for asset-driven creation and background generation tasks; Wan 3.0 Prime is a separate accelerated version, better suited to projects where generation wait time is the primary consideration. When choosing, test actual task completion times and final output quality. Do not interpret Prime's speed positioning as a fixed duration for the standard version, and do not judge image quality solely by the version name.

Getting Started

Assign a Role to Each Asset

In media, submit assets using types such as first_frame, last_frame, or reference_image/video/audio, clearly stating which ones constrain the subject and which provide motion or rhythm. The total array can contain up to 10 items.

Configure a Wan 3 Video Request

Explicitly select model=wan3.0-video for /wan/videos, and set duration, ratio, and resolution according to the task; audio is disabled by default, so set audio=true when sound output is needed.

Deliver Asynchronously and Validate

After saving task_id, retrieve the final video through /wan/tasks or a callback; check the actual duration, audio and video, and asset references, then save frames for the next shot.

Suggested trial: multi-asset product shot

Inputs and goal

The image defines the backpack's appearance, the video provides the walking motion, and the audio provides the rhythm; a person carries the backpack through a park, with the camera smoothly following and the motion coordinated with the rhythm.

Acceptance and next steps

Use media to distinguish reference_image, reference_video, and reference_audio; for output with sound, explicitly set audio=true, and check whether the roles of the assets are correctly expressed.

Usage boundaries

  • First and last frames cannot be mixed with reference assets in the same request. Choose frame mode when you need to lock the starting and ending images, and choose reference mode when you need to draw on appearance, motion, or sound; the total number of media items is limited to 10, and assets should be organized around the same creative goal to avoid stacking incompatible control methods together.
  • The reference video here is generation material and does not mean precise editing or seamless extension of the original footage. If the goal is to keep every existing shot unchanged and modify only local content, a dedicated editing workflow should be used; this model is better suited to recreating video clips based on prompts and references.
  • Successful asynchronous submission does not mean the video has already been completed. The application needs to save the task_id, check the task's final status, and retrieve the video link after success. file and link are asset input types, and you cannot assume that any file format or web page can be used directly based on this; validate the actual assets before integration.

Frequently Asked Questions

Does wan3.0-video require changing the model ID based on the creation method?

No. When calling /wan/videos, use model=wan3.0-video. Text-based creation is expressed through prompt, while first/last-frame or reference-based creation is expressed through media, and the system identifies the mode accordingly. First/last frames and reference materials should be submitted separately; do not mix them in the same request just to cover more control conditions.

What is the difference between duration=-1 and specifying a number of seconds?

You can specify an integer duration of 2–30 seconds, which is suitable for tasks with an existing editing rhythm or a clearly defined delivery length. duration=-1 lets the model intelligently choose the duration, making it suitable for initially exploring shot expression. If the video must fit a fixed placement or timeline, prioritize specifying the number of seconds to make subsequent editing easier to arrange.

Are reference audio and generated sound the same thing?

No. reference_audio is used to submit audio reference material, while audio controls whether the generated video includes sound, with a default of false. If you want the final video to include sound, explicitly enable audio and describe the sound goal in the prompt; submitting audio reference alone does not mean sound output has been enabled.

How do I retrieve a video generated by an asynchronous task?

After setting async=true, first save the returned task_id, then query the final task status through /wan/tasks. Successful videos are stored on this platform's CDN, and the result provides a video link and may include size and thumbnail information. The workflow should distinguish between submission, processing, and successful delivery, rather than only checking whether the request succeeded.

How does a reference video affect billing?

Reference videos are billed based on the combined actual seconds of input video and successfully generated output video; text, image, and audio inputs do not add video seconds. Therefore, for the same output length, usage charges may differ between creations with a reference video and text-only creations. View current pricing in Pricing; final usage is based on actual usage.