Image-and-text generation and video extension model for short-shot creation
kling-v1-6 is a video generation model in Kuaishou's Kling V1 series, suitable for turning text concepts or static images into short shots, then continuing creation from existing clips. It provides text-to-video, image-to-video, and video extension; pro mode can also use first and last frames to constrain the start and end of the visuals. When characters need to speak, photos and existing audio can be combined for animation and lip-syncing.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Video creation methods
Text-to-video, image-to-video, existing video extension
Generation duration
This platform's video entry point: 5 seconds or 10 seconds
Generation modes
std / pro; 4k mode is not supported
Video aspect ratios
16:9、9:16、1:1
First and last frame control
Image-to-video requires a first frame; pro can use a last frame, while std does not support last frames
Audio method
Native audio is not supported; talking photos use external audio-driven animation
Task delivery
Video link, video ID, task ID, and status; asynchronous processing and callbacks supported
The above durations, modes, and operation scope correspond to this platform's kling-v1-6 entry point. Photo lip-syncing is a combined creation feature and is not equivalent to the model's native audio.
Core Capabilities
From Text Concepts to Image Animation
When there is no ready-made image, use prompts to describe the subject, environment, and actions to generate a shot; when product images, character images, or scene images already exist, use the first frame to fix the starting point, then describe the desired changes. These two approaches are suitable for concept exploration and asset re-creation respectively, allowing short videos to begin with clear visual goals.
Use First and Last Frames to Set Shot Start and End
Image-to-video in pro mode can specify both the first and last frames to arrange changes in character poses, before-and-after product display states, or scene transitions. The last frame must work together with the first frame and cannot be used alone. It constrains the starting and ending images of the shot, not a frame-by-frame trajectory, and cannot replace dedicated camera movement control parameters.
Continue Creating from Existing Clips
Video extension can be used in std and pro modes: retain the generated result's video_id, then submit prompts for the next actions or plot. This lets you establish short shots first and then continue with subsequent content, without having to rebuild the scene from text each time. After the task is complete, retrieve the result through the video link for further editing.
Use Cases
Turn Static Product Assets into Short Videos
Input a product first-frame image, describe the display actions, background atmosphere, and visual changes, and generate short footage suitable for editing. Portrait format can be used for mobile displays, while landscape format can be used for presentation visuals; if there is already a clear final composition, add a last frame in pro, then supplement the finished video with subtitles, music, and brand information.
Storyboard Previsualization and Plot Continuation
Write a storyboard as a prompt with a clear subject, actions, and environment, generate a single shot first, then use video extension to try subsequent plot developments. The deliverables are watchable, editable video clips, suitable for discussing pacing and visual direction before formal filming or production, rather than directly replacing a complete multi-shot editing project.
Photo Voiceovers and Character Explanations
Prepare a clear front-facing photo of one person and recorded audio, select kling-v1-6 in the talking photo feature, and add action or expression requirements. A single task completes photo animation and audio lip-syncing, returning the final talking-head video and intermediate animation clips, suitable for short explanations, character greetings, or content previews.
How to choose this model
When to choose V1.6 versus pro
If the task is image animation, short-shot generation, or continuation of an existing clip, V1.6 offers more direct operation combinations; choose pro when you need to control the ending frame. Compared with kling-v1, you cannot carry over the older version's camera movement parameters to V1.6: kling-v1 camera movement has a 5-second condition, while V1.6 does not support camera_control, so choose according to actual control needs.
When to switch to later models
If you need to generate visuals and sound simultaneously, consider kling-v2-6 pro with audio support; if you need flexible durations of 3–15 seconds or 4k mode, consider kling-v3. V1.6 is better suited to short-video workflows that clearly use 5 or 10 seconds, with audio produced separately. External audio lip sync and native audio are different tasks and should not be mixed.
Get started
Prepare a text brief or first frame
Choose text2video to describe the scene, or choose image2video and provide start_image_url; write requirements for preserving people, products, or backgrounds into the prompt.
Specify the mode and final frame
Specify model=kling-v1-6 to /kling/videos, start with 5 seconds, and choose std/pro based on visual needs. Only pro image-to-video can use a final frame; std uses only the first frame.
Save assets after checking motion
Save the task_id asynchronously and obtain the video through task queries or callbacks; check the subject, final frame, and motion continuity. This model does not generate audio, so handle voiceover and music in post-production when needed; retain the video ID when continuing extensions.
Trial suggestion: Pro final frame for a product shot
Input and goal
The first frame shows a closed gift box, the final frame shows an open gift box, the product position inside the box is clear, the camera and tabletop remain consistent, and the box opens slowly.
Acceptance and next steps
First and final frames work together only under pro; check the geometry of the box lid and product. Do not request native audio for regular generation.
Usage Limits
kling-v1-6 does not support generate_audio, camera_control, or 4k mode. You can describe camera intent in the prompt, but this is not equivalent to parameterized camera movement; if the final video must have sound, you need to add voiceover and music separately, or use the photo lip-sync feature with audio.
std image-to-video cannot specify an end frame; pro end frames must also be submitted together with a start frame. Start and end frames are suitable for expressing beginning and ending compositions, but do not guarantee that all intermediate actions will occur along the intended path. Before production, clearly specify subject changes and action goals, then review the generated clip.
For talking photos, it is recommended to use a clear front-facing image of a single person. The driving audio must be an accessible link and supports mp3, wav, m4a, and aac, with a maximum size of 5MB. The audio length is recommended not to exceed the selected video duration; avoid putting long spoken passages directly into short video tasks.
Frequently Asked Questions
How can I ensure kling-v1-6 is used when making a request?
Explicitly specify model=kling-v1-6 in video generation or talking photo requests, rather than relying on the default selection after omission. For standard videos, you must also select text2video, image2video, or extend, and provide a prompt, start-frame image, or existing video ID according to the task.
How long a video can V1.6 generate?
The video generation interface supports 5-second or 10-second options, and talking photos also offer these two durations. To continue creating, you can use extend to submit an existing video_id and a follow-up prompt; extension continues an existing clip and does not mean generating a video of any length in a single request.
Why can't std mode use an end frame?
kling-v1-6 std does not support end-frame constraints. When a specific ending image is required, use pro image2video and provide both start_image_url and end_image_url; submitting only an end frame cannot complete this type of task, as a start frame is still required input.
Does a talking photo mean V1.6 can generate sound?
No. V1.6 does not support native audio. Talking photos use the audio you provide: the photo is animated first, then lip-synced to the audio. You need to submit image_url and audio_url; the prompt is mainly for actions or expressions during the animation stage, not for generating speech content.
How should generation tasks be integrated into an application?
You can use async=true to obtain a task_id and then query the task status, or configure callback_url to receive the completed result. Applications should distinguish between a submitted task and a completed video, and read video_url after completion; if you need to extend it later, you should also save video_id and cannot use task_id as a substitute.