Controlled short-form video creation driven by multi-image and video references
kling-v3-omni is Kuaishou Kling's video model for reference-based creation, also known as Kling 3.0 Omni. It can generate short videos from text or a first-frame image, and can also create and edit using multiple images, reference videos, or existing footage. Its value lies in bringing character, scene, and style assets into a single task, making it suitable for short-form video production with existing visual assets and a need for repeated adjustments to shot content.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Creation methods
Text-to-video, image-to-video, multi-image reference, video reference, and existing video editing
Generation duration
This platform's entry point: whole-number durations from 3–15 seconds
Aspect ratios
16:9, 9:16, 1:1
Generation modes
Standard generation: std / pro / 4k; Omni reference requests do not use 4k
Number of references
Up to 7 images without video; up to 4 images and 1 video with video
Audio controls
Standard generation can enable sound; video references can retain original audio and do not generate new audio simultaneously
Task output
Video link, video ID, task ID, duration, and status; asynchronous processing supported
The values and operating ranges above apply to this platform's access point; standard generation and Omni reference editing use different parameter combinations.
Core Capabilities
Turn visual assets into a creative foundation
Multi-image references let characters, scenes, and styles avoid being fully restated in text. You can place subject photos, environment images, and style images in the same task, then explicitly reference them in the prompt and explain their respective uses. This is especially suitable for creating with existing brand assets, allowing shots to develop around clear visual references instead of starting from scratch each time.
Distinguish between editing footage and reference footage
Video references offer two different uses: base uses a video as the foundation to be edited, allowing changes to elements, composition, color, weather, or overall style; feature extracts characteristics such as style and camera movement to guide a new video. Before choosing, clarify whether you want to alter the original footage or borrow its characteristics, which helps organize assets and prompts.
Organize motion from first and last frames
Image-to-video uses the first frame as a starting point and can also add a last-frame constraint for the ending image, making it suitable for short videos with existing storyboards or key visuals. Standard generation can select aspect ratio, duration, and whether to generate accompanying audio; when combining reference assets, switch to the Omni workflow and use assets and text together to describe subject actions and scene changes.
Use Cases
Product and brand short videos
Enter product images, brand environment images, and shot descriptions to generate short videos for social content or advertising concepts. Prompts should describe the product's position, display actions, and lighting atmosphere, while clearly specifying the role of each reference image. After delivery, focus on checking the product's appearance and image details, then adjust different creative versions around the same set of assets.
Creative revisions of existing footage
Use an existing video as base input and describe the style, background elements, or weather you want to change, such as transforming live-action footage into anime visuals. This is suitable for creating creative variations based on the original footage rather than rebuilding an entire storyboard. Audio can be retained or removed, with music, subtitles, and final editing added after the visual editing is complete.
Reference-driven storyboard exploration
Set a reference video as feature and pair it with character or scene images to explore new videos with specified visual styles or shot characteristics. This is suitable for director previews, storyboards, and series content concepts. For each task, first focus on a single shot objective; after delivering the short video, compare composition, action, and style before deciding the creative direction for subsequent shots.
How to choose this model
How to choose between it and kling-v3
If tasks mainly rely on text, first and last frames, and explicit camera parameter controls, kling-v3 is better suited to this workflow; if you need multi-image combinations, reference videos, or direct editing of existing shots, prioritize kling-v3-omni. Both can generate short videos with sound, but Omni does not support camera_control, so do not confuse reference camera movement with camera parameter controls.
How to choose between it and kling-o1
kling-o1 also supports all-in-one references and is not limited to processing a single image. Specific reasons to choose kling-v3-omni are the need for more flexible generation durations, sound in standard generation, or processing longer reference videos. Existing O1 workflows can retain similar asset organization approaches, but when switching models, reset the duration and audio combination rather than merely replacing the name.
Get started
Assign purposes to images and videos
Use text2video and image_list/video_list for reference creation, and specify the subject, style, and actions by number in the prompt. base is for editing existing videos, while feature is for video feature reference.
Set parameters by reference task
Specify model=kling-v3-omni for /kling/videos, starting with std/pro and whole-number 5-second increments. When reference video is included, set generate_audio=false; original audio is controlled by keep_original_sound. Do not mix reference tasks with standard generation parameters.
Check whether references are applied correctly
Save the task_id and video ID, and retrieve the final video through queries or callbacks; compare it with the assets to check the subject, actions, and preserved areas. If you choose to preserve the original audio, verify that track; when generating from scratch with sound enabled, then check the new audio.
Trial suggestion: character and camera-movement reference combination
Input and goal
Reference images define the character's clothing and appearance, while a reference video defines a slow panning shot; generate a clip of the character standing in a bookstore while maintaining the subject's characteristics.
Acceptance and next steps
Use image_list/video_list and reference the assets in the prompt; when video is included, do not generate new sound at the same time, and use keep_original_sound to preserve the original audio.
Usage Limits
Omni reference requests cannot use 4k, negative_prompt, cfg_scale, or camera_control. When visual adjustments are needed, express them through reference materials and positive prompts; regular generation may optionally use 4k, but this does not mean reference video editing can use the same mode.
When a reference video is included, generate_audio must be false; retaining the original audio is controlled by keep_original_sound. Base video editing can no longer specify first and last frames. The last frame for image-to-video must also be used together with a first frame; you cannot upload only an ending image.
Reference videos are limited to MP4/MOV, within 200MB, 24–60fps, and 3–15.5 seconds, and must also meet dimension and total pixel constraints. Materials must be referenced in the prompt; mixing first and last frames with reference images affects material order, so organize them consistently before submitting.
Frequently Asked Questions
What is the relationship between Kling O3 and kling-v3-omni?
Kling O3 and Kling 3.0 Omni are names for this model; kling-v3-omni is used when calling it on this platform. It corresponds to a different model from kling-v3 and kling-o1 respectively. When creating a task, explicitly specify the target model to avoid relying on the default selection.
How should I choose between multi-image reference and image-to-video?
If you already have a definite starting image, use image2video and provide a first frame, adding a last frame if necessary; if you need to combine character, environment, and style materials, use text2video with image_list. Reference images must not only be uploaded, but also explicitly referenced in the prompt, explaining their purpose in the shot.
Can I edit a video and also use a reference video to generate a new one?
Yes, but the materials have different roles. refer_type=base means editing this base video; feature means drawing on characteristics such as its style and camera movement to guide creation. The former is suitable for modifying existing content, while the latter is suitable for reference-driven new shots; the two task types should use different prompt objectives.
Can 4K and audio be used together in all tasks?
Not all tasks can be covered by the same combination. Regular generation supports 4k and audio options, while Omni reference requests do not use 4k; after adding a reference video, new audio generation must be disabled. If you want to retain the material's sound, set keep_original_sound instead of enabling generate_audio.
How do I obtain the generated video after submitting a task?
Submit the model, operation, and creative input to POST /kling/videos. Set async=true to first obtain a task_id, then query the task; you can also use callback_url to receive the result. After completion, obtain the video through video_url, and save the task ID, status, and material configuration for convenient tracking and iteration.