A single ivory folded-paper glider gently rises along a short smooth arc from the supplied first frame to the supplied final frame in three seconds. Maintain precisely one glider and unchanged folded geometry, canyon layers, camera position and soft lighting. Controlled miniature stop-motion feel, no wobbling wings, no strings, no cuts or additional objects.
Kling 3.0 Standard Image to Video API
kwaivgi/kling-v3.0-std/image-to-videoKling 3.0 Standard Image to Video animates still images at 720p, with native audio and optional end-frame guidance. Use a start frame and named subject references to guide appearance as the camera and subject move.
437/2,500
![Image Urls[0]](https://cdn.vidgo.ai/apis/models/kwaivgi/kling-v3.0-std/image-to-video/v1/01/input.png)
Name each element, add 2–4 JPG/PNG images or one MP4/MOV video, then reference it with @element_name.
Examples
REST API
Quick Start
Submit a Kling 3.0 Standard Image to Video request and use its task identifier to retrieve the video.
Authentication
Send your Vidgo API key in the Authorization header as a Bearer token.
- Submit endpoint
- POST
https://api.vidgo.ai/api/generate/submit - Authentication
- Authorization: Bearer YOUR_API_KEY
Submit Request
Send the model identifier and input object to the submission endpoint. Image examples contain a sample start frame; replace it with your own image URL.
REQUEST_BODY=$(cat <<'JSON'
{
"model": "kwaivgi/kling-v3.0-std/image-to-video",
"input": {
"prompt": "The camera slowly moves toward the subject. Soft daylight illuminates the scene as a gentle breeze moves through it.",
"duration": 5,
"multi_shots": false,
"sound": true,
"image_urls": [
"https://cdn.vidgo.ai/apis/models/kwaivgi/kling-v3.0-std/image-to-video/v1/01/input.png"
]
}
}
JSON
)
RESPONSE=$(curl --silent --show-error --fail-with-body \
--request POST \
--url "https://api.vidgo.ai/api/generate/submit" \
--header "Authorization: Bearer $VIDGO_API_KEY" \
--header "Content-Type: application/json" \
--data "$REQUEST_BODY")
CODE=$(printf '%s' "$RESPONSE" | jq -r '.code // empty')
if [ "$CODE" != "0" ] && [ "$CODE" != "200" ]; then
printf 'API error: %s
' "$RESPONSE" >&2
exit 1
fi
printf '%s
' "$RESPONSE"Query Task Status
Replace {task_id} in the status URL with the identifier returned on submission.
Status endpoint
https://api.vidgo.ai/api/generate/status/{task_id}Query this URL with the returned task_id. Continue until finished or failed; a completed video is listed in data.files.
not_startedrunningfinishedfailed{
"code": 200,
"data": {
"task_id": "example-task-id",
"status": "not_started",
"created_time": "2026-09-27T00:00:00Z"
}
}{
"code": 200,
"data": {
"task_id": "example-task-id",
"status": "finished",
"created_time": "2026-09-27T00:00:00Z",
"files": [
{
"file_type": "video",
"file_url": "https://example.com/generated-video.mp4"
}
]
}
}Complete Example
When status is finished, read file_url from video entries in files. Task IDs and output URLs below illustrate the response format.
set -euo pipefail
: "${VIDGO_API_KEY:?Set VIDGO_API_KEY in your environment}"
REQUEST_BODY=$(cat <<'JSON'
{
"model": "kwaivgi/kling-v3.0-std/image-to-video",
"input": {
"prompt": "The camera slowly moves toward the subject. Soft daylight illuminates the scene as a gentle breeze moves through it.",
"duration": 5,
"multi_shots": false,
"sound": true,
"image_urls": [
"https://cdn.vidgo.ai/apis/models/kwaivgi/kling-v3.0-std/image-to-video/v1/01/input.png"
]
}
}
JSON
)
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
--request POST \
--url "https://api.vidgo.ai/api/generate/submit" \
--header "Authorization: Bearer $VIDGO_API_KEY" \
--header "Content-Type: application/json" \
--data "$REQUEST_BODY")
TASK_ID=$(printf '%s' "$SUBMIT_RESPONSE" | jq -r '.data.task_id // .task_id // empty')
BUSINESS_CODE=$(printf '%s' "$SUBMIT_RESPONSE" | jq -r '.code // empty')
if [ "$BUSINESS_CODE" != "0" ] && [ "$BUSINESS_CODE" != "200" ]; then
printf 'Submit failed:
%s
' "$SUBMIT_RESPONSE" >&2
exit 1
fi
if [ -z "$TASK_ID" ]; then
printf 'Submit response did not include task_id:
%s
' "$SUBMIT_RESPONSE" >&2
exit 1
fi
START_TIME=$(date +%s)
POLL_DELAY=2
while true; do
if [ $(( $(date +%s) - START_TIME )) -ge 600 ]; then
printf 'Timed out after 600 seconds
' >&2
exit 1
fi
STATUS_RESPONSE=$(curl --silent --show-error --fail-with-body \
--url "https://api.vidgo.ai/api/generate/status/$TASK_ID" \
--header "Authorization: Bearer $VIDGO_API_KEY")
STATUS=$(printf '%s' "$STATUS_RESPONSE" | jq -r '.data.status // .status // empty')
BUSINESS_CODE=$(printf '%s' "$STATUS_RESPONSE" | jq -r '.code // empty')
if [ "$BUSINESS_CODE" != "0" ] && [ "$BUSINESS_CODE" != "200" ]; then
printf 'Status request failed:
%s
' "$STATUS_RESPONSE" >&2
exit 1
fi
case "$STATUS" in
finished)
printf '%s' "$STATUS_RESPONSE" | jq -r '(.data.files // .files // [])[]?.file_url'
break
;;
failed)
printf '%s' "$STATUS_RESPONSE" | jq -r '.data.error_message // .error_message // "Generation failed"' >&2
exit 1
;;
not_started|running)
sleep "$POLL_DELAY"
if [ "$POLL_DELAY" -lt 10 ]; then POLL_DELAY=$((POLL_DELAY + 1)); fi
;;
*)
printf 'Unexpected task status: %s
' "$STATUS" >&2
exit 1
;;
esac
doneRequest Parameters
Choose a single prompt or a sequence of shots, then set duration and sound for this endpoint.
| Field | Type | Requirement | Default | Usage |
|---|---|---|---|---|
| image_urls | string[] | Required | Set explicitly | Provide the first frame at index 0 and an optional last frame at index 1. Use public HTTP(S) image URLs. Use 1–2 frames for a single shot, or one start frame for multiple shots. |
| prompt | string | When multi_shots=false | Set explicitly | Describe movement from the reference frame in 1–2,500 characters. For multiple shots, write each scene in multi_prompt. |
| multi_shots | boolean | Optional | false | Choose false for one prompt, or true for separately timed shots with multi_prompt and sound=true. |
| multi_prompt | object[] | When multi_shots=true | Set explicitly | Add at least one object with prompt and duration. Each prompt uses 1–2,500 characters; each shot lasts 1–12 integer seconds. Shot times add up to duration. |
| duration | integer | Required | Set explicitly | Set the total clip length to an integer from 3 to 15 seconds. For multiple shots, use the sum of their durations. |
| sound | boolean | Optional | true | Use true for generated audio. A silent single shot uses false. Multiple shots use true. |
| aspect_ratio | string | Optional | From start frame | The first frame sets the output aspect ratio; prepare it in the composition you want. |
| kling_elements | object[] | Optional | Omit when unused | Provide reference element objects together with image_urls and refer to named elements in the prompt with @element_name. Each element has a name and either 2–4 JPG/PNG images up to 10 MB each or one MP4/MOV video up to 50 MB. An optional description can identify visual traits. |
Response Fields
Submission returns a task identifier. Query its status to retrieve generated video URLs.
| Field | Type | Usage |
|---|---|---|
| code | integer | Response code: 0 or 200 indicates a successful API operation. |
| data.task_id | string | Identifier returned on submission; use it to query the same task. |
| data.status | string | One of not_started, running, finished or failed. A successful submission is followed by status checks. |
| data.created_time | string | Task creation time returned by the service. |
| data.files | array | Generated media entries returned when the task finishes. |
| data.files[].file_type | string | Use entries with file_type=video for the generated clip. |
| data.files[].file_url | string | Video URL for playback or download. |
| data.error_message | string | null | Read the task error detail when status=failed. |
Task Lifecycle
Keep the returned task_id and follow status until finished or failed.
not_startedThe request is accepted and waiting to start. Keep its task_id for the next status check.
runningThe video is being generated. Continue checking the same task.
finishedGeneration is complete. Read the video entries in files.
failedGeneration ended with an error. Review error_message before submitting an adjusted request.
Error Handling
- Check timing before submittingFor multiple shots, use sound=true and ensure individual durations sum to a total of 3–15 seconds.
- Keep reference URLs accessibleUse public HTTP(S) media URLs that remain reachable while the task is processed.
- Resume interrupted status checksIf a status request fails, retry the status check using the same task_id. Submit a new generation only when you intend to create another clip.
- Read the returned errorFor a failed task, inspect error_message and adjust the relevant input before trying again.
Specifications
| Specification | Value | Usage |
|---|---|---|
| Input mode | Image to Video | A start frame plus motion direction; an end frame is optional. |
| Resolution | 720p | 720p video output for this tier. |
| Clip duration | 3–15 seconds | Whole seconds; each shot in a sequence uses 1–12 seconds. |
| Audio | sound | Generate audio with true; single-shot silent generation uses false. |
| Model identifier | kwaivgi/kling-v3.0-std/image-to-video | Public model value for this request. |
Related Models
Kling 3.0 Standard Image to Video API frequently asked questions
What is the Kling 3.0 Standard Image to Video API?
Kling 3.0 Standard Image to Video is a Kuaishou model for animating still images. It creates 720p footage with audio and frame guidance. Image conditioning guides the starting composition, while element references give recurring subjects a visual anchor. You can call it programmatically or try it from the playground above.
How do end frames guide Kling 3.0 Standard Image to Video?
Place the opening image at image_urls[0] and the desired closing image at image_urls[1]. Use one start frame for multiple shots; use a start frame and optional end frame for a single shot. Describe the action linking the two compositions.
How does Kling 3.0 Standard Image to Video use subject references?
Provide image_urls together with kling_elements, then use @element_name in the prompt to identify a recurring subject. Give each element a name and either 2–4 reference images or one reference video.
How are multi-shot scenes timed in Kling 3.0 Standard Image to Video?
Enable multi_shots and sound, then supply a prompt and an integer duration of 1–12 seconds for each shot. Set duration to the sum of the shot times, between 3 and 15 seconds. Use one start frame for multiple shots; use a start frame and optional end frame for a single shot.
How do I add audio in Kling 3.0 Standard Image to Video?
Set sound=true and describe the desired dialogue, ambient sound or effects alongside the visual action. Single shots can use sound=false for silent footage; multi-shot sequences use sound=true.
How does the first frame shape Kling 3.0 Standard Image to Video?
The first image establishes the starting composition and output aspect ratio. Crop it to the composition you want before submitting, then describe the subject action and camera movement that should develop from that frame.
When should I choose Kling 3.0 Standard over Kling 3.0 Pro?
Choose Standard for 720p scene exploration at a lower per-second rate. Choose Pro when the output needs 1080p detail; both offer audio and individually timed shots.
How much does a five-second Kling 3.0 Standard animation cost?
A silent single shot costs $0.675 at $0.135 per second. With sound enabled it costs $0.975 at $0.195 per second. Multiple shots use the audio rate and the sum of their durations.















