All models

doubao-seedance-2-5-260628 ★

ByteDanceVideo
Get your API key
doubao-seedance-2-5-260628

Create long takes and adapt videos with multimodal references

Seedance 2.5 is ByteDance's multimodal video generation model. doubao-seedance-2-5-260628 brings text creation, image-driven generation, audio and video references, video editing, and extension into a single workflow. It is suitable for projects that need more complete action sequences, multiple reference assets, or adaptations of existing videos, and can generate sound or silent videos up to 30 seconds long and up to 1080p.

ByteDanceModel brand
VideoModel type
Multimodal referenceCreation method
STANDARD APIs · QUICK SETUP

Bring this model into your workflow

Submit requests to the public API at api.acedata.cloud using the documented parameters, then use the results in your application.

API host
api.acedata.cloud
model
doubao-seedance-2-5-260628
Get your API key

Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.

Specifications and API features

Duration
4–30 seconds; duration:-1 for automatic duration
Resolution
480p, 720p, 1080p
Reference assets
Up to 30 reference images, 10 reference videos, and 10 reference audio clips, for a total of up to 50
Creation modes
Text, image-driven, multimodal reference; supports audio-only references, video editing, and extension
Aspect ratios
16:9, 9:16, 1:1, 4:3, 3:4, 21:9, and adaptive
Output and audio
MP4 or MOV; generate_audio controls audio output and is disabled by default
Task control
auto, reference, edit, extend; supports asynchronous tasks and completion callbacks

The above are the applicable specifications for this model on this platform. When calling it, select the corresponding settings for generation, editing, or extension mode.

Core Capabilities

Give each reference asset its own role

Images, videos, and audio can all participate in creation, or only audio references can be provided. When organizing assets, clearly specify in the text which images define the subject, which videos are used for actions or camera movement, and which audio is used for sound and rhythm, so that multiple references work together toward the same goal rather than simply stacking assets.

Move from generation to editing and continuation

In addition to generating videos from text or images, 2.5 also provides edit and extend modes, allowing modification or extension requests based on existing videos. This is suitable for clearly stating separately “what to preserve” and “what to change,” such as preserving the subject and camera movement, changing the atmosphere of the sky, or continuing to describe subsequent actions.

Organize visuals and sound in a single task

A maximum duration of 30 seconds is suitable for arranging an opening, main action, and conclusion, reducing the need to split a complete concept into too many short segments. When generate_audio is enabled, sound can be generated at the same time; when paired with audio references, you should still clearly describe the visual actions and sound intent to facilitate joint review of rhythm and expression.

Use Cases

Product demonstrations and advertising shots

Input product images, usage steps, and shot descriptions, use reference images to define the subject, and organize continuous segments around presentation, use, and ending shots. Portrait orientation is suitable for short-video ads, while landscape orientation is suitable for demonstration content; after delivering the video, focus on checking whether the product form and actions meet requirements, then create subtitles and brand information.

Character clips and sound-driven creation

Input character images, action videos, or audio references, describe the scene, action sequence, and sound atmosphere for the character, and generate video segments suitable for storyboard previews or content planning. If the idea is mainly driven by sound, you can also provide only audio references and text descriptions without first creating image or video assets.

Adaptation and supplementation of existing shots

Use an existing video as reference_video and clearly specify the elements to modify or the actions that need to continue. edit is used to modify existing content, and extend is used to lengthen clips; choose MP4 or MOV as the delivery format, and obtain results through asynchronous tasks for inclusion in subsequent editing and asset management workflows.

How to Choose This Model

Choose 2.5 When You Need Complete Sequences

Compared with the 4–15 second Seedance 2.0 series, 2.5 extends duration to 4–30 seconds and adds pure audio references as well as dedicated editing and extension modes. If the task involves longer action sequences, coordination of multiple assets, or adaptation of existing videos, 2.5 is a better fit; existing short-shot workflows can retain the original model for comparison.

Balance Duration and Clarity

2.5 supports up to 1080p, so it should not be understood as exceeding 2.0 in every specification. Consider 2.0 Standard when 4k output is needed; for short-video iteration requiring only 480p or 720p, compare 2.0 Fast or Mini. When choosing, first assess whether the complete action is usable and whether the assets are effective, then determine whether the output clarity meets delivery requirements.

Getting Started

Organize content and Assets

In content, assign specific roles to images, videos, and audio, and explain the purpose of each asset in the text; organize first/last-frame constraints and subject references separately according to the task.

Select the Correct Version and Shot Settings

Specify model=doubao-seedance-2-5-260628 for /seedance/videos; start by testing with duration=10, resolution=1080p, and a clear aspect ratio. When sound is needed, explicitly set generate_audio=true; it is disabled by default. For editing/extension, separately select omni_reference_task_type and provide reference_video; do not reuse the fixed aspect ratio of standard generation.

Save the Task and Final Frame

For asynchronous requests, first obtain task_id, then query /seedance/tasks or receive a callback; after completion, check the subject, actions, ending, and audio, then save the final video and returned final frame as needed.

Trial Recommendation: Complete Action Sequences and Continuation

Input and Goal

Start with product and character references: a person packs a backpack, leaves the room, and walks down a corridor into a sunlit courtyard, all in one continuous shot while maintaining subject and audio continuity.

Acceptance Criteria and Next Steps

Use 20 seconds to clearly show three actions, check the transitions before considering extend; extension or editing requires reference_video, and choose adaptive and duration according to each mode.

Usage Limits

  • Both editing and extending must provide reference_video. edit must use ratio:adaptive and duration:-1; extend also requires adaptive, with duration set to 4–30 seconds or automatic. Do not directly apply fixed aspect-ratio settings for regular generation to video modification tasks.
  • reference mode requires at least one image, video, or audio reference. The maximum number of assets is not a creative goal; character, action, and sound instructions should have clear roles. A mismatch between input assets and task type will cause failure, and modification requirements should be organized around a clear creative objective.
  • Each text item can be up to 1000 characters. Complex scripts need to distill the main subject, action, camera, and sound priorities. Image addresses must be placed in the url object of image_url and cannot be entered directly as strings; reference audio and video should also use their respective object structures and corresponding roles.

Frequently Asked Questions

Can Seedance 2.5 generate 4k videos?

This model supports 480p, 720p, and 1080p, with 1080p as the maximum. If delivery explicitly requires 4k, consider Seedance 2.0 Standard; if the main need is longer clips, audio-only references, editing, or extending, then 2.5 is more suitable, so there is no need to choose based only on version numbers.

Can I upload only audio as a reference?

Yes. 2.5 supports using reference_audio alone, without also providing images or videos. Pair it with text describing the visual subject, action, and scene; using audio references in creation does not automatically enable audio output, so set generate_audio:true when sound is needed.

How do I distinguish between editing a video and extending a video?

Choose edit to modify existing content and extend to continue expanding a clip; both require a reference video and an adaptive aspect ratio. The duration for edit must be -1; extend can be set to 4–30 seconds or -1. Prompts should respectively emphasize what needs to change or the action that needs to continue.

What is the difference between first/last frame images and subject references?

first_frame and last_frame are used to specify the starting or ending image, while reference_image provides reference information such as the subject. Use the first frame when you want a video to begin moving from an image; use a subject reference and explain its purpose in the text when you want to create a new scene or action based on an image.

How do I submit a task and get the video?

Submit model and content to /seedance/videos, with model set to doubao-seedance-2-5-260628. Set async:true to first obtain a task_id, then query the task status; you can also set callback_url to receive a completion notification, and finally obtain the finished video from the video link in the result.