All models

kling-v3-turbo

KuaishouVideo
Get your API key
Kling 3.0 Turbo · Video Generation

Generate short videos with native audio from text or a first-frame image

Use kling-v3-turbo to generate 3–15-second video clips. Choose std 720p or pro 1080p output, and structure prompts around the subject, action, scene, and sound.

STANDARD APIs · QUICK SETUP

Bring this model into your workflow

Submit requests to the public API at api.acedata.cloud using the documented parameters, then use the results in your application.

API host
api.acedata.cloud
model
kling-v3-turbo
Get your API key

Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.

Capabilities

Generate shots from text descriptions

Use action=text2video, explicitly specify model=kling-v3-turbo, and describe the subject, action, composition, and sound requirements in the prompt.

Start from a first-frame image

Use action=image2video and start_image_url to make the input image the starting point of the shot, then describe the subject's movement and scene changes.

Native audio generated with the visuals

This model includes native audio and does not provide an option to disable it. Omit generate_audio or set it to true; after generation, you should still verify that the sound meets delivery requirements.

Core Specifications

Invocation ID
kling-v3-turbo
Endpoint
POST /kling/videos
Generation Methods
Text-to-video, first-frame image-to-video
Duration
3–15 whole seconds
Quality
std: 720p; pro: 1080p
Audio
Native audio, with no option to disable it

Use Cases

Individual shots in short films

Fit one clearly defined action into 3–15 seconds, validate the visuals and sound first, then combine them into a longer work.

Animate images

Use a person, product, or scene image as the first frame, and describe changes around the subject already present in the image.

Sound-enabled creative samples

Compare image quality, motion, and audio using real business prompts before deciding whether to use the generated result. The API does not guarantee that a single generation is ready for direct delivery.

How to Choose a Model

Choose Turbo for its input and audio combination

When a task requires only text or a single first-frame image and accepts native audio, you can start with kling-v3-turbo. std and pro correspond to different output resolutions; see Pricing for actual billing.

First check the specialized capabilities of other models

When you need an end frame, camera movement targets, independent negative prompts, or cfg_scale, do not apply these parameters directly to Turbo. For multi-image references and editing existing videos, choose the corresponding specialized model and operation, and verify against the documentation for that endpoint.

Getting Started

Prepare text or a single first frame

For text-to-video, use text2video and prompt; for image-to-video, use image2video and start_image_url. Write the subject's actions, camera work, and desired sounds into the description; do not submit an end frame.

Choose Turbo output settings

Explicitly set model=kling-v3-turbo for /kling/videos, choose an integer between 3–15 seconds, with std at 720p and pro at 1080p. Audio is generated natively; omit generate_audio or set it to true.

Review visuals and sound together

Use async or callback_url to retrieve completed tasks, and save the task_id and video ID; check motion, character or product details, and listen to sound effects and dialogue. For muted delivery, mute the finished video in post-production.

Trial suggestion: a first-frame ad with native sound

Input and goal

Use an image of a soda cup as the first frame, with ice cubes gently falling into the cup, bubbles rising, a fixed camera, and a quiet background, emphasizing crisp sounds.

Review and next steps

Choose 5 seconds and std or pro, then check the cup, ice cube motion, and audio; this model has no audio-off switch and does not accept an end frame.

Usage limits

  • End frames, camera_control, standalone negative_prompt, and cfg_scale are not supported. Content to avoid should be written in the prompt.
  • std outputs 720p and pro outputs 1080p; do not promise 4K configurations from other models for this model.
  • Native audio does not provide an off switch; do not pass generate_audio=false. Muted videos require post-processing.
  • Duration must be an integer between 3–15 seconds. Evaluate output quality and generation time using actual samples; do not infer speed guarantees from the model name.

Frequently Asked Questions

How do I start text-to-video?

Submit action=text2video, model=kling-v3-turbo, and prompt to POST /kling/videos, and set a valid quality and integer duration.

What is required for first-frame image-to-video?

Use action=image2video and provide start_image_url. This model supports first-frame guidance but does not support end_image_url end frames.

Which durations and resolutions can I choose?

Durations are integer values from 3–15 seconds; std is 720p and pro is 1080p.

Can I turn off generated audio?

This model includes native audio and does not provide an off switch. Omit generate_audio or set it to true; for muted output, post-process the completed video.

Are camera controls or standalone negative prompts supported?

camera_control, negative_prompt, and cfg_scale are not supported. Describe camera requirements and content to avoid in the prompt.

How do I retrieve asynchronous task results?

After setting async=true, first obtain the task_id, then query status and final video results through the task API; you can also use callback_url to receive completion notifications.