Preserve the adult performer, red costume, two silks and empty circus tent from the image. Securely supported in the hip wrap, she slowly rotates a quarter turn; free silk tails follow with realistic weight. A gentle camera arc follows without cutting. Keep natural anatomy and finish in a stable held pose. Audio: soft breathing, silk friction and quiet tent ambience; no music or speech. No text, logos or watermarks.
Kling O3 Standard Image to Video API
kwaivgi/kling-video-o3-std/image-to-videoGenerate 3–15-second videos with Kling O3 Standard from a starting image and an optional ending frame. Supports single-shot and multi-shot generation. Explicitly set sound; multi-shot requires audio.
410/2,500
![Image Urls[0]](https://cdn.vidgo.ai/apis/models/kwaivgi/kling-video-o3-std/image-to-video/v1/03/input-start.png)
![Image Urls[1] (optional)](https://cdn.vidgo.ai/apis/models/kwaivgi/kling-video-o3-std/image-to-video/v1/03/input-end.png)
Examples
REST API Reference
Quick Start
Submit a task and query its status.
Step 1: Set up authentication
Create an API key in the dashboard and attach Authorization: Bearer <API_KEY> when submitting a task.
- Submit Endpoint
- POST
https://api.vidgo.ai/api/generate/submit - Authorization Header
- Authorization: Bearer VIDGO_API_KEY
Step 2: Submit a task
POST /api/generate/submit: kwaivgi/kling-video-o3-std/image-to-video
REQUEST_BODY=$(cat <<'JSON'
{
"model": "kwaivgi/kling-video-o3-std/image-to-video",
"input": {
"duration": 3,
"sound": false,
"multi_shots": false,
"prompt": "Fixed-camera miniature scene. Move from the supplied raised-drawbridge start frame to the lowered-drawbridge end frame. The single rigid wooden bridge slowly rotates down around its bottom hinge at the doorway until it reaches the opposite bank; both chains extend naturally. Preserve the towers, moat, tabletop, lighting and framing. No people, new structures or camera movement. No text, logos or watermarks.",
"image_urls": [
"https://cdn.vidgo.ai/apis/models/kwaivgi/kling-video-o3-std/image-to-video/v1/03/input-start.png",
"https://cdn.vidgo.ai/apis/models/kwaivgi/kling-video-o3-std/image-to-video/v1/03/input-end.png"
]
}
}
JSON
)
RESPONSE=$(curl --silent --show-error --fail-with-body \
--request POST \
--url "https://api.vidgo.ai/api/generate/submit" \
--header "Authorization: Bearer $VIDGO_API_KEY" \
--header "Content-Type: application/json" \
--data "$REQUEST_BODY")
CODE=$(printf '%s' "$RESPONSE" | jq -r '.code // empty')
if [ "$CODE" != "0" ] && [ "$CODE" != "200" ]; then
printf 'API error: %s
' "$RESPONSE" >&2
exit 1
fi
printf '%s
' "$RESPONSE"Step 3: Poll for completion
Poll with task_id while status is not_started or running, and stop at finished or failed. On success, read data.files[].file_url; on failure, read data.error_message.
Status Endpoint
GET https://api.vidgo.ai/api/generate/status/{task_id}Poll with task_id while status is not_started or running, and stop at finished or failed. On success, read data.files[].file_url; on failure, read data.error_message.
not_startedrunningfinishedfailed{
"code": 200,
"data": {
"task_id": "task-submitted-example",
"status": "not_started",
"created_time": "2026-09-22T00:00:00Z"
}
}{
"code": 200,
"data": {
"task_id": "97JVN8DM9N5UJLYO",
"status": "finished",
"files": [
{
"file_type": "video",
"file_url": "https://cdn.vidgo.ai/apis/models/kwaivgi/kling-video-o3-std/image-to-video/v1/03/output.mp4"
}
],
"created_time": "2026-09-22T18:58:33",
"error_message": null,
"progress": 100
}
}Complete executable script
Expand to review an end-to-end script with automatic polling, error handling, and timeout safeguards.
set -euo pipefail
: "${VIDGO_API_KEY:?Set VIDGO_API_KEY in your environment}"
REQUEST_BODY=$(cat <<'JSON'
{
"model": "kwaivgi/kling-video-o3-std/image-to-video",
"input": {
"duration": 3,
"sound": false,
"multi_shots": false,
"prompt": "Fixed-camera miniature scene. Move from the supplied raised-drawbridge start frame to the lowered-drawbridge end frame. The single rigid wooden bridge slowly rotates down around its bottom hinge at the doorway until it reaches the opposite bank; both chains extend naturally. Preserve the towers, moat, tabletop, lighting and framing. No people, new structures or camera movement. No text, logos or watermarks.",
"image_urls": [
"https://cdn.vidgo.ai/apis/models/kwaivgi/kling-video-o3-std/image-to-video/v1/03/input-start.png",
"https://cdn.vidgo.ai/apis/models/kwaivgi/kling-video-o3-std/image-to-video/v1/03/input-end.png"
]
}
}
JSON
)
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
--request POST \
--url "https://api.vidgo.ai/api/generate/submit" \
--header "Authorization: Bearer $VIDGO_API_KEY" \
--header "Content-Type: application/json" \
--data "$REQUEST_BODY")
TASK_ID=$(printf '%s' "$SUBMIT_RESPONSE" | jq -r '.data.task_id // .task_id // empty')
BUSINESS_CODE=$(printf '%s' "$SUBMIT_RESPONSE" | jq -r '.code // empty')
if [ "$BUSINESS_CODE" != "0" ] && [ "$BUSINESS_CODE" != "200" ]; then
printf 'Submit failed:
%s
' "$SUBMIT_RESPONSE" >&2
exit 1
fi
if [ -z "$TASK_ID" ]; then
printf 'Submit response did not include task_id:
%s
' "$SUBMIT_RESPONSE" >&2
exit 1
fi
START_TIME=$(date +%s)
POLL_DELAY=2
while true; do
if [ $(( $(date +%s) - START_TIME )) -ge 600 ]; then
printf 'Timed out after 600 seconds
' >&2
exit 1
fi
STATUS_RESPONSE=$(curl --silent --show-error --fail-with-body \
--url "https://api.vidgo.ai/api/generate/status/$TASK_ID" \
--header "Authorization: Bearer $VIDGO_API_KEY")
STATUS=$(printf '%s' "$STATUS_RESPONSE" | jq -r '.data.status // .status // empty')
BUSINESS_CODE=$(printf '%s' "$STATUS_RESPONSE" | jq -r '.code // empty')
if [ "$BUSINESS_CODE" != "0" ] && [ "$BUSINESS_CODE" != "200" ]; then
printf 'Status request failed:
%s
' "$STATUS_RESPONSE" >&2
exit 1
fi
case "$STATUS" in
finished)
printf '%s' "$STATUS_RESPONSE" | jq -r '(.data.files // .files // [])[]?.file_url'
break
;;
failed)
printf '%s' "$STATUS_RESPONSE" | jq -r '.data.error_message // .error_message // "Generation failed"' >&2
exit 1
;;
not_started|running)
sleep "$POLL_DELAY"
if [ "$POLL_DELAY" -lt 10 ]; then POLL_DELAY=$((POLL_DELAY + 1)); fi
;;
*)
printf 'Unexpected task status: %s
' "$STATUS" >&2
exit 1
;;
esac
doneRequest Parameters (input object)
Put generation parameters inside input. Use standard JSON numbers and booleans. For compatibility, integer strings such as "5" are accepted. Boolean strings true/1/yes/y/on mean true; false/0/no/n/off mean false. These strings are case-insensitive and trimmed. Numeric 1 and 0 are also accepted for boolean fields. Numeric and boolean prompt values are converted to text; 0, false and null are treated as empty. Objects and arrays are not accepted as prompts. The API validates prompts up to 2,500 characters. Some multi-shot generation requests have failed when a shot prompt exceeded 512 characters. We recommend keeping each shot prompt within 512 characters; this recommendation does not change the API validation limit.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
| prompt | string | Conditional | - | Required when multi_shots=false. Describe the scene, action and camera movement. Use nonblank text, up to 2,500 characters after trimming. Omit prompt when multi_shots=true. |
| multi_shots | boolean | Yes | - | Required: explicitly send false for a single shot or true for multiple shots. Single-shot mode requires prompt. Multi-shot mode requires multi_prompt and sound=true, with no nonblank top-level prompt. Omitting this field is an error. |
| multi_prompt | array | Conditional | - | Required when multi_shots=true; omit in single-shot mode. Supply at least one shot, each with a nonblank prompt of up to 2,500 characters and an integer duration of 1–12 seconds. Shot durations must sum to the top-level duration (3–15 seconds). Extra fields inside a shot are ignored. |
| duration | integer | Yes | - | Required integer from 3 to 15 seconds. In multi-shot mode, this must equal the sum of all shot durations. Credits are calculated using this value. |
| sound | boolean | Yes | - | Required: explicitly send true to generate audio or false for video without audio. Single-shot mode accepts either value; multi-shot mode requires true. |
| aspect_ratio | string | No | - | May be omitted. If provided, use 16:9, 9:16 or 1:1. Image-to-video ignores this value; framing is determined by the input image. |
| image_urls | array | Yes | - | Required: an array of 1–2 images. The first is the starting frame; the optional second is the ending frame. Accepts public HTTP(S) URLs, image Data URIs or raw Base64 image data. |
Response Fields (Status Query)
Details returned by GET /api/generate/status/{task_id}:
| Field | Type | Description |
|---|---|---|
| code | integer | HTTP/business response status code (200 indicates success). |
| data.task_id | string | Globally unique task identifier. |
| data.status | string | Task lifecycle state: not_started, running, finished, or failed. |
| data.files | array | Array of output assets containing file_url and file_type upon completion. |
| data.error_message | string | null | Error diagnostic details if the task status is failed. |
Task Lifecycle
Clients should poll status until reaching either the finished or failed terminal state:
not_startedQueued
runningGenerating
finishedReady
failedFailed
Polling & Error Handling
- Polling frequencyStart polling with a 2 to 3-second interval, gradually increasing to 5 seconds for extended takes.
- Network resiliencyTransient 5xx responses or timeouts do not signify task failure; retry status requests after a short backoff.
- Webhook callbacksProvide a top-level callback_url in your submission payload to receive completion notifications automatically.
Specifications
| Specification | Value | Description |
|---|---|---|
| Model | kwaivgi/kling-video-o3-std/image-to-video | Kuaishou Kling O3 unified multimodal image-to-video standard tier supporting single-frame initiation and dual-keyframe 720p physical motion interpolation. |
| Duration | 3-15 s | Required integer, 3–15 seconds. In multi-shot mode it must equal the sum of shot durations. Billing uses this duration. |
Kling O3 Standard Image to Video
Generate 3–15-second videos with Kling O3 Standard from a starting image and an optional ending frame. Supports single-shot and multi-shot generation. Explicitly set sound; multi-shot requires audio.
Why Choose This Model
Dual-Keyframe ControlUpload distinct start and end frames, allowing the model to interpolate physically plausible motion paths to control opening and closing compositions.
Strong Identity RetentionFaithfully preserves facial features, garment textures, and background perspective from source stills without drift or structural distortion.
Native Action Audio SynthesisSynthesizes ambient room tone and motion-timed Foley sound effects natively from visual semantics without external audio post-processing.
High-Efficiency 720p OutputDelivers smooth 720p animation at an accessible 10 credits per second, providing ideal throughput for high-volume asset workflows.
Multi-Shot & Duration FreedomSupports continuous clips between 3 and 15 seconds with multi-shot narrative sequencing, expanding still artwork into dynamic short films.
Parameters
| Parameter | Requirement | Description |
|---|---|---|
| prompt | Conditional | Required when multi_shots=false. Describe the scene, action and camera movement. Use nonblank text, up to 2,500 characters after trimming. Omit prompt when multi_shots=true. Default - |
| multi_shots | Yes | Required: explicitly send false for a single shot or true for multiple shots. Single-shot mode requires prompt. Multi-shot mode requires multi_prompt and sound=true, with no nonblank top-level prompt. Omitting this field is an error. Default - |
| multi_prompt | Conditional | Required when multi_shots=true; omit in single-shot mode. Supply at least one shot, each with a nonblank prompt of up to 2,500 characters and an integer duration of 1–12 seconds. Shot durations must sum to the top-level duration (3–15 seconds). Extra fields inside a shot are ignored. Default - |
| duration | Yes | Required integer from 3 to 15 seconds. In multi-shot mode, this must equal the sum of all shot durations. Credits are calculated using this value. Default - |
| sound | Yes | Required: explicitly send true to generate audio or false for video without audio. Single-shot mode accepts either value; multi-shot mode requires true. Default - |
| aspect_ratio | No | May be omitted. If provided, use 16:9, 9:16 or 1:1. Image-to-video ignores this value; framing is determined by the input image. Default - |
| image_urls | Yes | Required: an array of 1–2 images. The first is the starting frame; the optional second is the ending frame. Accepts public HTTP(S) URLs, image Data URIs or raw Base64 image data. Default - |
How to Use
Upload key visual assetsProvide 1 initial frame image (required), and optionally upload 1 end-frame image to anchor the closing composition.
Guide motion and cameraWrite a descriptive motion prompt detailing character actions, environmental dynamics, and desired camera trajectory.
Select runtime and audioChoose an integer duration between 3 and 15 seconds, then toggle synchronized audio on or off.
Submit generation taskDispatch the request via API or click Run; the engine calculates spatio-temporal physics between keyframes.
Preview and export videoReview the rendered 720p animation to inspect kinetic fluidity and audio-visual synchronization, then download the MP4 file.
Pricing
Credits = top-level duration × per-second rate. 1 credit = $0.005. If generation fails, consumed credits are automatically and fully refunded.
| Usage | Rate | Details |
|---|---|---|
| Without sound | 10 credits/s ($0.050/s) | Same rate for text-to-video and image-to-video. |
| With sound | 13 credits/s ($0.065/s) | Same rate for text-to-video and image-to-video. |
Best Use Cases
Still Photography AnimationTransform static portraiture or landscape photography into vibrant 720p motion vignettes for digital portfolios.
E-commerce Product Stills to MotionAnimate static catalog photographs with fabric drapes, fluid splashes, and subtle camera movements.
Keyframe Guided TransitionsInput wide-angle and close-up stills as start and end frames to interpolate seamless camera pushes.
Concept Art and Illustration MotionBring 2D illustrations and character designs to life with blinking, hair flutter, and dynamic lighting.
Pro Tips
- When uploading 2 images, ensure lighting angles and subject scale remain logically related to facilitate natural physical interpolation.
- Focus prompts on transitional motion paths, describing how the subject moves from the opening posture to the ending state.
- Use clean, uncompressed JPEG or PNG source images with the primary subject centered in the frame.
- For talking character animations, upload a portrait with a neutral mouth posture and include spoken lines in quotation marks.
- Respect natural physics and momentum; avoid prompting sudden, instantaneous posture flips that conflict with the initial still.
Notes
- image_urls supports 1 to 2 publicly accessible HTTP(S) URLs or Base64 image payloads (first image as start frame, second as end frame).
- When multi_shots is enabled, leave the top-level prompt empty, define each scene in multi_prompt, and set sound=true.
- Generation is asynchronous; monitor task progress with the returned task_id until status reaches finished or failed.
Kling O3 Standard Image to Video API FAQ
What does this endpoint generate?
Generate 3–15-second videos with Kling O3 Standard from a starting image and an optional ending frame. Supports single-shot and multi-shot generation. Explicitly set sound; multi-shot requires audio.
How do I use starting and ending frames?
Required: an array of 1–2 images. The first is the starting frame; the optional second is the ending frame. Accepts public HTTP(S) URLs, image Data URIs or raw Base64 image data.
How is the video aspect ratio determined?
May be omitted. If provided, use 16:9, 9:16 or 1:1. Image-to-video ignores this value; framing is determined by the input image.
How do I set up multiple shots?
Required: explicitly send false for a single shot or true for multiple shots. Single-shot mode requires prompt. Multi-shot mode requires multi_prompt and sound=true, with no nonblank top-level prompt. Omitting this field is an error. Required when multi_shots=true; omit in single-shot mode. Supply at least one shot, each with a nonblank prompt of up to 2,500 characters and an integer duration of 1–12 seconds. Shot durations must sum to the top-level duration (3–15 seconds). Extra fields inside a shot are ignored.
How do I control audio and prompt length?
Required: explicitly send true to generate audio or false for video without audio. Single-shot mode accepts either value; multi-shot mode requires true. The API validates prompts up to 2,500 characters. Some multi-shot generation requests have failed when a shot prompt exceeded 512 characters. We recommend keeping each shot prompt within 512 characters; this recommendation does not change the API validation limit.
How are duration and credits calculated?
Duration is 3–15 seconds. Video without audio costs 10 credits/s; video with audio costs 13 credits/s. Credits equal the top-level duration multiplied by the applicable rate.
What happens if generation fails or my request times out?
Invalid parameters are rejected before a generation task is created or credits are deducted. If an accepted generation task later reaches failed, its deducted credits are refunded. A client timeout alone does not mean the task failed; query its task_id before submitting again.