All models

doubao-seedance-2-0-260128 ★

ByteDanceVideo
Get your API key
doubao-seedance-2-0-260128

Standard Multimodal Reference Video Version for High-Definition Final Output

doubao-seedance-2-0-260128 is ByteDance Seedance 2.0 Standard Edition, suitable for combining text, character images, and audio/video references into short videos. It can bring static images to life while using reference materials to constrain subjects, actions, and camera rhythm, with support for sound generation. When you need high-definition output, character references, and short-shot creation, this is a version worth prioritizing.

ByteDanceModel brand
VideoModel type
Multimodal referenceCreation method
STANDARD APIs · QUICK SETUP

Bring this model into your workflow

Submit requests to the public API at api.acedata.cloud using the documented parameters, then use the results in your application.

API host
api.acedata.cloud
model
doubao-seedance-2-0-260128
Get your API key

Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.

Specifications and API Features

Creation method
Text-to-video, image-to-video, first-and-last-frame generation, multimodal reference
Output resolution
480p, 720p, 1080p, 4k
Generation duration
4–15 seconds; duration=-1 for automatic duration
Aspect ratio
16:9, 4:3, 1:1, 3:4, 9:16, 21:9, adaptive
Reference materials
Up to 9 images; up to 3 audio files and 3 videos each
Sound control
Enable generate_audio for sound generation; disabled by default
Delivery method
Video link; supports asynchronous tasks, callbacks, and last-frame return

The above describes this platform's supported invocation scope for this model. Native capabilities and material requirements for each creation mode should be understood separately.

Core Capabilities

Separate control of subject references and scene creation

Use reference_image to provide a person or subject, then specify the scene, clothing, actions, and camera work with text, allowing the reference image to provide identity and appearance cues rather than fixing the entire composition. If you want the video to begin moving from the original image composition, use first_frame instead, which is suitable for turning portraits, product images, or illustrations into dynamic shots.

Express actions and rhythm with audio and video

In addition to images, you can add reference_video and reference_audio to provide cues for actions, camera movement, sound, and rhythm. Text should specify which parts of each material need to be referenced, avoiding conflicting requirements from different materials; when you need a final video with sound, enable generate_audio to bring visuals and audio into the same generation process.

From preview formats to high-definition delivery

The Standard Edition offers resolution options from 480p to 4k and covers landscape, portrait, square, and ultra-wide formats. You can first validate the subject and actions, then choose a higher resolution for delivery shots; use return_last_frame to retrieve the final frame, making it convenient to preview the ending or prepare visual material for the next creation.

Applicable Scenarios

Short Product Showcase Shots

Input product images and shot descriptions, such as a slow push-in, lateral movement, or showcasing material details, to generate standalone clips for advertising edits. Use First Frame mode when the starting layout needs to be preserved, use Reference mode when presenting the same subject in a new scene, then choose landscape, portrait, or square aspect ratio based on the placement.

Cross-Scene Character Shorts

Provide clear reference images of people you own or are authorized to use, describe the character's actions indoors, on the street, or in natural environments, and create multiple short shots of the same character. Reference images handle appearance cues, while text handles scenes and actions; suitable for character concept presentations and storyboard sequences. Before delivery, check each segment to ensure the character's appearance and actions meet requirements.

Action and Sound Proofs of Concept

Combine action example videos, rhythmic audio, and scene prompts to explore visual expression for dance, sports, or scenes with ambient sound. First clarify whether you want to reference the action, shot, or sound, generate short clips for the creative team to discuss, then download the videos for the editing workflow. Avoid treating every element in a reference asset as a target that must be copied.

How to Choose This Model

Standard Version vs. Fast and Mini

When 1080p or 4k output is needed, prioritize this Standard version; 2.0 Fast and Mini support 480p and 720p resolutions, making them better suited for previews and lightweight creation. All three can use character and audio-video references, but they are different model variants. Do not assume that all parameter combinations and final output performance are exactly the same just by replacing the name.

HD Short Shots vs. Longer Editing Tasks

This model is suitable for 4–15-second HD short shots and multimodal reference generation. If you need up to 30 seconds, audio-only references, more assets, or editing and extending existing videos, choose Seedance 2.5. The key trade-off between the two is HD output and workflow type, rather than interpreting the newer version as having higher specifications in every respect.

Getting Started

Organize content and assets

In content, assign specific roles to images, videos, and audio, and describe the purpose of each asset in the text; organize first-and-last-frame constraints and subject references separately by task.

Select the correct version and shot settings

Specify model=doubao-seedance-2-0-260128 for /seedance/videos; first test with duration=5, resolution=720p, and a clear aspect ratio. When sound is needed, explicitly set generate_audio=true; it is disabled by default.

Save tasks and final frames

For asynchronous requests, first obtain the task_id, then query /seedance/tasks or receive a callback; after completion, check the subject, motion, ending, and audio, then save the finished video and returned final frame as needed.

Trial suggestion: high-definition product motion reference

Input and objective

A product image defines a pair of athletic shoes, and a reference video defines a simple walking motion. Generate a shoe showcase shot on an outdoor track, preserving the shoe shape and color scheme, with no text.

Review and next steps

First clarify the roles of reference_image and reference_video, then check the motion and product details; 4k is only for combinations supported by this standard version.

Usage boundaries

  • First-frame, first-and-last-frame, and full-modal reference are mutually exclusive modes. first_frame and last_frame cannot be mixed with various reference tags; in multimodal creation, text can describe the intended use of images as start or end frames, but when the first and last frames need to be locked, use the dedicated first-and-last-frame mode instead.
  • Reference audio must be wav or mp3, with each clip 2–15 seconds and no more than 15 MB; reference video must be mp4 or mov, with each clip also 2–15 seconds. The total duration of audio and video respectively cannot exceed 15 seconds. Trim assets before preparing them to avoid failures during processing.
  • This model is for generation; editing, extension, web-connected tools, or mov output selection should not be considered its features. People and character assets must be owned or authorized, and photos should be clear, front-facing, and unobstructed; reference assets are used to guide generation and do not mean that every motion and detail will be reproduced exactly.

Frequently Asked Questions

What is the difference between a reference image and a first-frame image?

reference_image is used to provide character or subject cues, while the scene and actions can still be redefined with text; first_frame makes the video start from the specified image. To place the same character in a new environment, choose a reference image; to make the composition of an existing photo come alive directly, choose a first frame. Do not mix the two tags.

How do I generate a video with sound?

Set generate_audio to true, and describe the desired sounds or ambient audio in the prompt to request a video with sound; this option is disabled by default. If reference audio is needed, you must use the audio_url object and the reference_audio tag, while also providing reference_image or reference_video; you cannot upload reference audio alone. Sound generation and audio reference are different controls; uploading audio cannot replace the setting for generating sound. When not using reference audio, you can also describe the desired sound in text.

Can this version generate 30-second videos or extend existing videos?

This version generates videos from 4–15 seconds, and duration=-1 can also be used to automatically determine the duration. For up to 30 seconds or editing and extending existing videos, please choose Seedance 2.5; adding reference_video is a reference-generation method and does not mean editing or continuing the original video.

How should image URLs be specified in requests?

Submit the exact model ID and content array to POST /seedance/videos. Image items use an image_url object, for example containing url internally, rather than writing the address directly as a string; use type, role, and text prompts to describe their purpose, and it is recommended to set resolution and duration using top-level fields.

How do I retrieve the completed video after submission?

You can set async=true to obtain task_id and then query the task, or provide callback_url to receive a completion notification. After receiving the completed result, read data.video_url to download the video; do not treat a successful submission as meaning the final video is already complete. If you need an ending image, you can also enable return_last_frame.