Kling Videos Generation API Integration Guide

This article will introduce a Kling Videos Generation API integration guide, which can generate Kling official videos by entering custom parameters.

Application Process

To use the Kling Videos Generation API, first go to the 胖狐中转 Console to obtain your API Token and keep it for later use.

If you have not logged in or registered yet, you will be automatically redirected to the login page and invited to register and log in. After completion, you will automatically return to the current page.

One API Token can call all services on the platform; there is no need to apply separately for each service. The first application comes with free credits for a free trial; when credits are insufficient, you can top up the general balance in the Console.

📘 Full documentation: Kling Videos Generation API →

Basic Usage

First, let's understand the basic usage method: by entering the prompt prompt, generation action action, first-frame reference image start_image_url, and model model, you can obtain the processed result. First, you need to simply pass an action field with the value text2video. It mainly includes three actions: text-to-video (text2video), image-to-video (image2video), and extend video (extend). Then we also need to enter the model model. Currently, the main models include kling-v1, kling-v1-6, kling-v2-master, kling-v2-1-master, kling-v2-5-turbo, kling-v2-6, kling-v3, kling-v3-omni, and kling-o1. The specific content is as follows:

You can see that we have set Request Headers here, including:

  • accept: The format of the response result you want to receive. Fill in application/json here, which is JSON format.
  • authorization: The key for calling the API. After applying, you can directly select it from the dropdown.

Additionally, Request Body is set, including:

  • model: The model for generating videos, mainly including kling-v1, kling-v1-6, kling-v2-master, kling-v2-1-master, kling-v2-5-turbo, kling-v2-6, kling-v3, kling-v3-omni, and kling-o1.
  • mode: The mode for generating videos. Optional values are standard mode std, fast mode pro, and native 4K mode 4k. Among them, 4k only supports kling-v3 and kling-v3-omni, and is incompatible with camera_control (camera movement control).
  • action: The action of this video generation task, mainly including three actions: text-to-video (text2video), image-to-video (image2video), and extend video (extend).
  • start_image_url: When selecting the image-to-video action image2video, the link to the first-frame reference image that must be uploaded.
  • end_image_url: Optional for image-to-video, specifies the last frame.
  • duration: Video duration, in seconds. kling-v3 and kling-v3-omni support integer durations from 3 to 15 seconds; kling-o1 only supports 5 seconds; other models support 5 or 10 seconds.
  • generate_audio: Whether to generate audio simultaneously, optional, Boolean value. Supported by kling-v3, kling-v3-omni, and kling-v2-6 (pro mode only). Defaults to false.
  • aspect_ratio: Video aspect ratio, optional, supports 16:9, 9:16, and 1:1, default is 16:9.
  • cfg_scale: Relevance strength, range [0,1]; the larger the value, the more closely it matches the prompt.
  • camera_control: Optional, object parameters for controlling camera movement, supporting type/simple presets and configurations such as horizontal, vertical, pan, tilt, roll, and zoom.
  • negative_prompt: Optional, negative prompts that you do not want to appear, up to 200 characters.
  • image_list: Omni reference image list, applicable to the kling-o1 and kling-v3-omni models. For usage, see "Omni Universal Reference" below.
  • video_list: Omni reference video list (supports video editing), applicable to the kling-o1 and kling-v3-omni models. For usage, see "Omni Universal Reference" below.
  • prompt: Prompt.
  • callback_url: The URL that needs callback results.
  • async: Optional. When set to true, the API immediately returns task_id; there is no need to provide callback_url, and then the result can be obtained by polling through the corresponding task query API.

After selection, you can find that the corresponding code is also generated on the right, as shown in the figure:

Click the "Try" button to test. As shown in the figure above, here we get the following result:

{
  "success": true,
  "video_id": "900798310464749610",
  "video_url": "https://cdn.acedata.cloud/assets/examples/kling/6c68c267-065b-4423-b66b-a0e4c59ee0d5-6a664a591a53.mp4",
  "duration": "5.041",
  "state": "succeed",
  "task_id": "6c68c267-065b-4423-b66b-a0e4c59ee0d5"
}

There are multiple fields in the returned result, introduced as follows:

  • success, the status of the video generation task at this time.
  • task_id, the ID of the video generation task at this time.
  • video_id, the video ID of the video generation task at this time.
  • video_url, the video link of the video generation task at this time.
  • duration, the video link duration of the video generation task at this time.
  • state, the status of the video generation task at this time.

You can see that we have obtained satisfactory video information. We only need to obtain the generated Kling video based on the video link address in data in the result.

Additionally, if you want to generate the corresponding integration code, you can directly copy the generated code. For example, the CURL code is as follows:

curl -X POST 'https://api.ace.324567.xyz/kling/videos' \
-H 'accept: application/json' \
-H 'authorization: Bearer {token}' \
-H 'content-type: application/json' \
-d '{
  "action": "text2video",
  "model": "kling-v3",
  "prompt": "White ceramic coffee mug on glossy marble countertop with morning window light. Camera slowly rotates 360 degrees around the mug, pausing briefly at the handle."
}'

Model Capability Matrix

Different models vary greatly in their support for parameters. The following matrix is compiled from the Kling official video models documentation. Before calling, please first verify whether the current model / mode / duration combination supports the functionality you need; otherwise, errors such as model/mode/duration(...) is not supported with image_tail will be returned.

Model Mode end_image_url (First and Last Frames) generate_audio (Soundtrack) camera_control (Camera Movement) Notes
kling-v1 std / pro ✅ Only duration=5 ❌ ✅ Only duration=5 extend does not support negative_prompt and cfg_scale
kling-v1-6 std ❌ ❌ ❌ Multi-image-to-video, extend available in all modes
kling-v1-6 pro ✅ ❌ ❌
kling-v2-master — ❌ ❌ ❌ Single mode, only duration=5/10
kling-v2-1-master — ❌ ❌ ❌ Single mode, only duration=5/10
kling-v2-5-turbo std ❌ ❌ ❌
kling-v2-5-turbo pro ✅ ❌ ❌
kling-v2-6 std ❌ ❌ ❌
kling-v2-6 pro ✅ ✅ ❌ The only non-v3 model that also supports soundtrack
kling-v3 std / pro ✅ ✅ ✅ duration range: 3–15 seconds
kling-v3 4k ✅ ✅ ❌ 4K mode is incompatible with camera movement
kling-v3-omni std / pro / 4k ✅ ✅ ❌
kling-o1 std / pro ✅ ❌ ❌ Only supports duration=5

Notes:

  • mode=4k is supported only by kling-v3 and kling-v3-omni; it is mutually exclusive with camera_control (camera movement).
  • end_image_url can only be used together with start_image_url when action=image2video. Passing only end_image_url (without start_image_url) will be rejected.
  • kling-v3 / kling-v3-omni accept any integer duration from 3–15 seconds; kling-o1 only accepts 5; all other models only accept 5 or 10.
  • generate_audio defaults to false. Only kling-v3, kling-v3-omni, and kling-v2-6 (pro mode) support it.

Video Extension Feature

If you want to continue generating an already generated Kling video, you can set the parameter action to extend, and input the ID of the video that needs to be continued. The video ID can be obtained based on the basic usage, as shown in the image below:

At this point, you can see that the video ID is:

"video_id": "030bb06d-98d4-4044-9042-0aa0822e8c8c"

Note that the video_id in this video is the ID of the generated video. If you do not know how to generate a video, you can refer to the basic usage above to generate a video.

Next, we must fill in the prompt for the next step of extension to customize the generated video, and can specify the following content:

  • model: The model for generating videos, mainly including kling-v1, kling-v1-5, and kling-v1-6 models.
  • mode: The mode for generating videos. Optional values are standard mode std, fast mode pro, and native 4K mode 4k (supported only by kling-v3 and kling-v3-omni, incompatible with camera movement control).
  • duration: The video duration for this video generation task, mainly including 5s and 10s.
  • start_image_url: When selecting the image-to-video behavior image2video, the link to the first-frame reference image that must be uploaded.
  • prompt: Prompt.

The filling example is as follows:

After filling it in, the following code is automatically generated:

The corresponding Python code:

import requests

url = "https://api.ace.324567.xyz/kling/videos"

headers = {
    "accept": "application/json",
    "authorization": "Bearer {token}",
    "content-type": "application/json"
}

payload = {
    "action": "extend",
    "model": "kling-v1",
    "video_id": "030bb06d-98d4-4044-9042-0aa0822e8c8c",
    "prompt": "White ceramic coffee mug on glossy marble countertop with morning window light. Camera slowly rotates 360 degrees around the mug, pausing briefly at the handle.",
    "duration": 10
}

response = requests.post(url, json=payload, headers=headers)
print(response.text)

Click Run, and you can find that you will get a result as follows:

{
  "success": true,
  "video_id": "bbc3b105-ac72-4de2-8390-0cb37dc7d41e",
  "video_url": "https://cdn.acedata.cloud/assets/examples/gemini/04a043bd-6b23-4b4e-945c-ce48158c3eee-3a89912507c7.mp4",
  "duration": "9.6",
  "state": "succeed",
  "task_id": "3ece87e6-3ee3-4f5e-bd70-5ae5eca89a23"
}

As you can see, the result content is consistent with the above, which implements the video extension feature.

Omni Universal Reference (Video Editing / Reference Video / Multi-Image Reference)

kling-o1 and kling-v3-omni are two independent models, both supporting the "universal reference" capability. Based on text-to-video (action=text2video), you can additionally pass in reference images or reference videos to achieve multi-image reference, reference videos, and direct editing of existing videos.

Core convention: Reference materials must be cited in prompt in the form of <<<image_1>>>, <<<video_1>>> (numbering starts from 1), referring to the material at the corresponding position in image_list / video_list, for the model to apply these references. If materials are passed without being cited in the prompt, the materials will be ignored.

Security note: The current API does not expose element_list. The IDs in the Kling Element Library are not tenant-isolated. Before an Element Management API that provides tenant isolation is available, please use image_list to pass subject reference images.

Omni requests do not support negative_prompt, cfg_scale, or camera_control, and cannot use mode=4k. When a reference video is included, generate_audio must be false.

Reference Video and Video Editing (video_list)

video_list is used to pass reference videos and is the most commonly used scenario for this capability. The array element fields are as follows:

  • video_url: Reference video link, cannot be empty. At most 1 MP4/MOV video, file size ≤200MB, frame rate 24–60fps. kling-o1 requires a duration of 3–10 seconds and both width and height of 700–2160px; kling-v3-omni requires a duration of 3–15.5 seconds, both width and height of 700–4553px, total pixels ≤8,294,400, and an aspect ratio of 0.4–2.
  • refer_type: Reference type, optional base (default, the base video to be edited, i.e., "directly edit the video", where elements can be added/removed/modified, composition changed, style changed, colors changed, weather changed, etc.) or feature (feature reference, referencing its style / camera movement / continuation for the next shot).
  • keep_original_sound: Whether to retain the original video audio, optional yes (retain) or no (remove).

Note: When a reference video exists, generate_audio must be false. A video with refer_type=base cannot additionally specify a first frame / last frame.

The CURL example for editing an existing video (changing the video to an anime style) is as follows:

curl -X POST 'https://api.ace.324567.xyz/kling/videos' \
-H 'accept: application/json' \
-H 'authorization: Bearer {token}' \
-H 'content-type: application/json' \
-d '{
  "action": "text2video",
  "model": "kling-o1",
  "mode": "std",
  "duration": 5,
  "prompt": "Change <<<video_1>>> into a cinematic anime style, while preserving the original motion and composition",
  "video_list": [
    {
      "video_url": "https://cdn.acedata.cloud/your-reference-video.mp4",
      "refer_type": "base",
      "keep_original_sound": "no"
    }
  ]
}'

Multi-image Reference (image_list)

image_list is used to pass reference images (elements / scenes / styles, etc.). The array element fields are as follows:

  • image_url: Reference image link, cannot be empty. Requirements: .jpg/.jpeg/.png format; file size ≤10MB; shortest side ≥300px; aspect ratio 1:2.5 ~ 2.5:1.
  • type: Optional. When not provided, it is used as a pure reference image; when first_frame / end_frame is provided, it is used as the first frame / last frame respectively (equivalent to start_image_url / end_image_url).

When using it, reference it in prompt as <<<image_1>>>, <<<image_2>>>. Quantity limits: reference images ≤ 7 when no reference video exists; reference images ≤ 4 when a reference video exists. When only passing the first / last frame, you can also directly use start_image_url / end_image_url, but the last frame must be used together with the first frame.

Note: If start_image_url / end_image_url and image_list are passed at the same time, the first / last frames will be placed before image_list, which may affect the index correspondence of <<<image_N>>>. It is recommended to choose one: when first / last frames are needed, specify them directly with type in image_list, and do not mix them with start_image_url / end_image_url.

CURL example for generating a video with multi-image reference:

curl -X POST 'https://api.ace.324567.xyz/kling/videos' \
-H 'accept: application/json' \
-H 'authorization: Bearer {token}' \
-H 'content-type: application/json' \
-d '{
  "action": "text2video",
  "model": "kling-o1",
  "mode": "std",
  "duration": 5,
  "prompt": "Make the character in <<<image_1>>> stand in the scene of <<<image_2>>>, with cinematic lighting",
  "image_list": [
    { "image_url": "https://cdn.acedata.cloud/subject.png" },
    { "image_url": "https://cdn.acedata.cloud/scene.png" }
  ]
}'

Asynchronous Callback

Since the Kling Videos Generation API takes a relatively long time to generate, approximately 1–2 minutes, if the API does not respond for a long time, the HTTP request will remain connected, causing additional system resource consumption. Therefore, this API also provides support for asynchronous callbacks.

The overall process is: when the client initiates a request, it additionally specifies a callback_url field. After the client initiates the API request, the API immediately returns a result containing a task_id field, representing the current task ID. When the task is completed, the generated video result will be sent in POST JSON format to the callback_url specified by the client, which also includes the task_id field, so that the task result can be associated through the ID.

Below, we will learn how to operate it specifically through an example.

First, a Webhook callback is a service that can receive HTTP requests. Developers should replace it with the URL of their own deployed HTTP server. For the convenience of demonstration, a public Webhook example website https://webhook.site/ is used here. Open this website to obtain a Webhook URL, as shown in the figure:

Copy this URL, and it can be used as a Webhook. The example here is https://webhook.site/624b2c78-6dbd-4618-9d2b-b32eade6d8c3.

Next, we can set the callback_url field to the above Webhook URL and fill in the corresponding parameters. The specific content is shown in the figure:

Click Run, and you can see that a result is immediately returned, as follows:

{
  "task_id": "20068983-0cc9-4c6a-aeb6-9c6a3c668be0"
}

After waiting for a moment, we can observe the generated video result at https://webhook.site/624b2c78-6dbd-4618-9d2b-b32eade6d8c3, as shown in the figure:

The content is as follows:

{
    "success": true,
    "video_id": "030bb06d-98d4-4044-9042-0aa0822e8c8c",
    "video_url": "https://cdn.acedata.cloud/assets/examples/gemini/04a043bd-6b23-4b4e-945c-ce48158c3eee-3a89912507c7.mp4",
    "duration": "5.1",
    "state": "succeed",
    "task_id": "20068983-0cc9-4c6a-aeb6-9c6a3c668be0"
}

You can see that there is a task_id field in the result. The other fields are similar to those above, and task association can be achieved through this field.

Error Handling

When calling the API, if an error occurs, the API will return the corresponding error code and information. For example:

  • 400 token_mismatched: Bad request, possibly due to missing or invalid parameters.
  • 400 api_not_implemented: Bad request, possibly due to missing or invalid parameters.
  • 401 invalid_token: Unauthorized, invalid or missing authorization token.
  • 429 too_many_requests: Too many requests, you have exceeded the rate limit.
  • 500 api_error: Internal server error, something went wrong on the server.

Error Response Example

{
  "success": false,
  "error": {
    "code": "api_error",
    "message": "fetch failed"
  },
  "trace_id": "2cf86e86-22a4-46e1-ac2f-032c0f2a4e89"
}

Conclusion

Through this document, you have learned how to use the Kling Videos Generation API to generate videos by entering prompts and first-frame reference images. We hope this document can help you better integrate with and use this API. If you have any questions, please feel free to contact our technical support team.

Kling 3.0 Turbo

model="kling-v3-turbo" supports text-to-video and first-frame image-to-video, with integer durations from 3–15 seconds. mode="std" outputs 720p, and mode="pro" outputs 1080p. This model includes native audio and does not provide an option to disable it; omit generate_audio or set it to true. This model does not support end frames, camera movement objects, independent negative prompts, or cfg_scale. Please directly describe the content you need or want to avoid in prompt.

Multi-Shot Videos

kling-v3 and kling-v3-omni support multi_shot=true. shot_type="intelligence" automatically creates shots based on prompt; shot_type="customize" provides 1–6 shots through multi_prompt, with each item containing an index that increases consecutively from 1, a prompt of up to 512 characters, and an integer duration of at least 1 second. The sum of all shot durations must equal the total duration. Custom shots do not use the global prompt. Multi-shot videos are billed according to the selected model, quality, audio configuration, and total duration.