Traditional Chinese ink-and-watercolor animation on textured rice paper. An empty narrow Jiangnan alley in rain, pale stone paving, dark tiled eaves. One fallen ochre oil-paper umbrella lies sideways in the foreground. A gentle gust rolls it half a turn through a shallow puddle; its reflection blooms into soft ink ripples. Locked low camera, continuous movement, restrained monochrome palette with the single ochre accent. No people. No text, logos or watermarks.
Kling O3 Standard Text to Video API
kwaivgi/kling-video-o3-std/text-to-videoGenerate 3–15-second videos from text prompts with Kling O3 Standard. Supports single-shot and multi-shot generation. Explicitly set sound; multi-shot requires audio.
Multi-shot requires sound. The form sets duration to the shot total.
Duration: 5 s (3-15)
Examples
REST API Reference
Quick Start
Submit a task and query its status.
Step 1: Set up authentication
Create an API key in the dashboard and attach Authorization: Bearer <API_KEY> when submitting a task.
- Submit Endpoint
- POST
https://api.vidgo.ai/api/generate/submit - Authorization Header
- Authorization: Bearer VIDGO_API_KEY
Step 2: Submit a task
POST /api/generate/submit: kwaivgi/kling-video-o3-std/text-to-video
REQUEST_BODY=$(cat <<'JSON'
{
"model": "kwaivgi/kling-video-o3-std/text-to-video",
"input": {
"duration": 5,
"sound": true,
"multi_shots": true,
"multi_prompt": [
{
"prompt": "Photoreal cinematic science-fiction inside a compact orbital maintenance module. Medium shot: one adult astronaut with short dark hair, wearing a plain off-white flight suit with orange elbow patches, gently catches a single small silver wrench floating in front of her chest. Her body and the wrench drift in zero gravity. Soft blue Earth light through a round window, warm cabin practicals. Quiet ventilation hum and fabric rustle. Controlled natural movement. No text, logos or watermarks.",
"duration": 2
},
{
"prompt": "Cut to a close view of the same astronaut's hand and orange elbow patch in the same orbital cabin. She places the same silver wrench onto a dark magnetic tool panel. The wrench attaches with one crisp metallic click; she releases it and her hand floats gently away. Maintain the same lighting, costume and wrench. End on the stationary attached wrench. Only quiet ventilation and the synchronized magnetic click; no music or speech. No text, logos or watermarks.",
"duration": 3
}
],
"aspect_ratio": "16:9"
}
}
JSON
)
RESPONSE=$(curl --silent --show-error --fail-with-body \
--request POST \
--url "https://api.vidgo.ai/api/generate/submit" \
--header "Authorization: Bearer $VIDGO_API_KEY" \
--header "Content-Type: application/json" \
--data "$REQUEST_BODY")
CODE=$(printf '%s' "$RESPONSE" | jq -r '.code // empty')
if [ "$CODE" != "0" ] && [ "$CODE" != "200" ]; then
printf 'API error: %s
' "$RESPONSE" >&2
exit 1
fi
printf '%s
' "$RESPONSE"Step 3: Poll for completion
Poll with task_id while status is not_started or running, and stop at finished or failed. On success, read data.files[].file_url; on failure, read data.error_message.
Status Endpoint
GET https://api.vidgo.ai/api/generate/status/{task_id}Poll with task_id while status is not_started or running, and stop at finished or failed. On success, read data.files[].file_url; on failure, read data.error_message.
not_startedrunningfinishedfailed{
"code": 200,
"data": {
"task_id": "task-submitted-example",
"status": "not_started",
"created_time": "2026-09-22T00:00:00Z"
}
}{
"code": 200,
"data": {
"task_id": "57TMAHB6XPG6VHL9",
"status": "finished",
"files": [
{
"file_type": "video",
"file_url": "https://cdn.vidgo.ai/apis/models/kwaivgi/kling-video-o3-std/text-to-video/v1/01/output.mp4"
}
],
"created_time": "2026-09-22T18:37:04",
"error_message": null,
"progress": 100
}
}Complete executable script
Expand to review an end-to-end script with automatic polling, error handling, and timeout safeguards.
set -euo pipefail
: "${VIDGO_API_KEY:?Set VIDGO_API_KEY in your environment}"
REQUEST_BODY=$(cat <<'JSON'
{
"model": "kwaivgi/kling-video-o3-std/text-to-video",
"input": {
"duration": 5,
"sound": true,
"multi_shots": true,
"multi_prompt": [
{
"prompt": "Photoreal cinematic science-fiction inside a compact orbital maintenance module. Medium shot: one adult astronaut with short dark hair, wearing a plain off-white flight suit with orange elbow patches, gently catches a single small silver wrench floating in front of her chest. Her body and the wrench drift in zero gravity. Soft blue Earth light through a round window, warm cabin practicals. Quiet ventilation hum and fabric rustle. Controlled natural movement. No text, logos or watermarks.",
"duration": 2
},
{
"prompt": "Cut to a close view of the same astronaut's hand and orange elbow patch in the same orbital cabin. She places the same silver wrench onto a dark magnetic tool panel. The wrench attaches with one crisp metallic click; she releases it and her hand floats gently away. Maintain the same lighting, costume and wrench. End on the stationary attached wrench. Only quiet ventilation and the synchronized magnetic click; no music or speech. No text, logos or watermarks.",
"duration": 3
}
],
"aspect_ratio": "16:9"
}
}
JSON
)
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
--request POST \
--url "https://api.vidgo.ai/api/generate/submit" \
--header "Authorization: Bearer $VIDGO_API_KEY" \
--header "Content-Type: application/json" \
--data "$REQUEST_BODY")
TASK_ID=$(printf '%s' "$SUBMIT_RESPONSE" | jq -r '.data.task_id // .task_id // empty')
BUSINESS_CODE=$(printf '%s' "$SUBMIT_RESPONSE" | jq -r '.code // empty')
if [ "$BUSINESS_CODE" != "0" ] && [ "$BUSINESS_CODE" != "200" ]; then
printf 'Submit failed:
%s
' "$SUBMIT_RESPONSE" >&2
exit 1
fi
if [ -z "$TASK_ID" ]; then
printf 'Submit response did not include task_id:
%s
' "$SUBMIT_RESPONSE" >&2
exit 1
fi
START_TIME=$(date +%s)
POLL_DELAY=2
while true; do
if [ $(( $(date +%s) - START_TIME )) -ge 600 ]; then
printf 'Timed out after 600 seconds
' >&2
exit 1
fi
STATUS_RESPONSE=$(curl --silent --show-error --fail-with-body \
--url "https://api.vidgo.ai/api/generate/status/$TASK_ID" \
--header "Authorization: Bearer $VIDGO_API_KEY")
STATUS=$(printf '%s' "$STATUS_RESPONSE" | jq -r '.data.status // .status // empty')
BUSINESS_CODE=$(printf '%s' "$STATUS_RESPONSE" | jq -r '.code // empty')
if [ "$BUSINESS_CODE" != "0" ] && [ "$BUSINESS_CODE" != "200" ]; then
printf 'Status request failed:
%s
' "$STATUS_RESPONSE" >&2
exit 1
fi
case "$STATUS" in
finished)
printf '%s' "$STATUS_RESPONSE" | jq -r '(.data.files // .files // [])[]?.file_url'
break
;;
failed)
printf '%s' "$STATUS_RESPONSE" | jq -r '.data.error_message // .error_message // "Generation failed"' >&2
exit 1
;;
not_started|running)
sleep "$POLL_DELAY"
if [ "$POLL_DELAY" -lt 10 ]; then POLL_DELAY=$((POLL_DELAY + 1)); fi
;;
*)
printf 'Unexpected task status: %s
' "$STATUS" >&2
exit 1
;;
esac
doneRequest Parameters (input object)
Put generation parameters inside input. Use standard JSON numbers and booleans. For compatibility, integer strings such as "5" are accepted. Boolean strings true/1/yes/y/on mean true; false/0/no/n/off mean false. These strings are case-insensitive and trimmed. Numeric 1 and 0 are also accepted for boolean fields. Numeric and boolean prompt values are converted to text; 0, false and null are treated as empty. Objects and arrays are not accepted as prompts. The API validates prompts up to 2,500 characters. Some multi-shot generation requests have failed when a shot prompt exceeded 512 characters. We recommend keeping each shot prompt within 512 characters; this recommendation does not change the API validation limit.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
| prompt | string | Conditional | - | Required when multi_shots=false. Describe the scene, action and camera movement. Use nonblank text, up to 2,500 characters after trimming. Omit prompt when multi_shots=true. |
| multi_shots | boolean | Yes | - | Required: explicitly send false for a single shot or true for multiple shots. Single-shot mode requires prompt. Multi-shot mode requires multi_prompt and sound=true, with no nonblank top-level prompt. Omitting this field is an error. |
| multi_prompt | array | Conditional | - | Required when multi_shots=true; omit in single-shot mode. Supply at least one shot, each with a nonblank prompt of up to 2,500 characters and an integer duration of 1–12 seconds. Shot durations must sum to the top-level duration (3–15 seconds). Extra fields inside a shot are ignored. |
| duration | integer | Yes | - | Required integer from 3 to 15 seconds. In multi-shot mode, this must equal the sum of all shot durations. Credits are calculated using this value. |
| sound | boolean | Yes | - | Required: explicitly send true to generate audio or false for video without audio. Single-shot mode accepts either value; multi-shot mode requires true. |
| aspect_ratio | string | No | - | Optional: 16:9 (landscape), 9:16 (portrait) or 1:1 (square). Sets the video aspect ratio. |
Response Fields (Status Query)
Details returned by GET /api/generate/status/{task_id}:
| Field | Type | Description |
|---|---|---|
| code | integer | HTTP/business response status code (200 indicates success). |
| data.task_id | string | Globally unique task identifier. |
| data.status | string | Task lifecycle state: not_started, running, finished, or failed. |
| data.files | array | Array of output assets containing file_url and file_type upon completion. |
| data.error_message | string | null | Error diagnostic details if the task status is failed. |
Task Lifecycle
Clients should poll status until reaching either the finished or failed terminal state:
not_startedQueued
runningGenerating
finishedReady
failedFailed
Polling & Error Handling
- Polling frequencyStart polling with a 2 to 3-second interval, gradually increasing to 5 seconds for extended takes.
- Network resiliencyTransient 5xx responses or timeouts do not signify task failure; retry status requests after a short backoff.
- Webhook callbacksProvide a top-level callback_url in your submission payload to receive completion notifications automatically.
Specifications
| Specification | Value | Description |
|---|---|---|
| Model | kwaivgi/kling-video-o3-std/text-to-video | Kuaishou Kling O3 unified multimodal text-to-video standard tier with native 720p resolution and multi-shot storyboarding. |
| Duration | 3-15 s | Required integer, 3–15 seconds. In multi-shot mode it must equal the sum of shot durations. Billing uses this duration. |
Kling O3 Standard Text to Video
Generate 3–15-second videos from text prompts with Kling O3 Standard. Supports single-shot and multi-shot generation. Explicitly set sound; multi-shot requires audio.
Why Choose This Model
Efficient 720p SynthesisBalances dependable visual fidelity with rapid turnaround cycles, offering ample throughput for creative prototyping and social content pipelines.
Multi-Shot StoryboardingOrchestrate multiple narrative beats in a single pass, configuring per-shot prompts and durations for cohesive chronological storytelling.
Native Audiovisual SyncGenerate synchronized ambient soundscapes, realistic Foley sound effects, and multilingual lip-sync in the same inference pass without post-dubbing.
Visual Chain-of-ThoughtLeverages unified multimodal reasoning to plan spatial composition, camera trajectories, and lighting dynamics before rendering frames.
Versatile Framing OptionsSwitch effortlessly between 16:9 widescreen, 9:16 vertical, and 1:1 square aspect ratios to fit desktop players and mobile feeds natively.
Parameters
| Parameter | Requirement | Description |
|---|---|---|
| prompt | Conditional | Required when multi_shots=false. Describe the scene, action and camera movement. Use nonblank text, up to 2,500 characters after trimming. Omit prompt when multi_shots=true. Default - |
| multi_shots | Yes | Required: explicitly send false for a single shot or true for multiple shots. Single-shot mode requires prompt. Multi-shot mode requires multi_prompt and sound=true, with no nonblank top-level prompt. Omitting this field is an error. Default - |
| multi_prompt | Conditional | Required when multi_shots=true; omit in single-shot mode. Supply at least one shot, each with a nonblank prompt of up to 2,500 characters and an integer duration of 1–12 seconds. Shot durations must sum to the top-level duration (3–15 seconds). Extra fields inside a shot are ignored. Default - |
| duration | Yes | Required integer from 3 to 15 seconds. In multi-shot mode, this must equal the sum of all shot durations. Credits are calculated using this value. Default - |
| sound | Yes | Required: explicitly send true to generate audio or false for video without audio. Single-shot mode accepts either value; multi-shot mode requires true. Default - |
| aspect_ratio | No | Optional: 16:9 (landscape), 9:16 (portrait) or 1:1 (square). Sets the video aspect ratio. Default - |
How to Use
Define subject and contextDetail character appearance, key kinetic actions, and environmental lighting in your descriptive text prompt.
Choose camera workflowSelect a single continuous take or activate multi-shot mode to sequence individual scene prompts and durations.
Specify length and framingPick an integer duration between 3 and 15 seconds alongside your target aspect ratio (16:9, 9:16, or 1:1).
Toggle synchronized audioEnable sound to synthesize matching environmental noise and speech; multi-shot mode enforces sound automatically.
Submit and retrieveDispatch the asynchronous job to receive a unique task ID, then monitor progress until the final MP4 asset is ready to preview and download.
Pricing
Credits = top-level duration × per-second rate. 1 credit = $0.005. If generation fails, consumed credits are automatically and fully refunded.
| Usage | Rate | Details |
|---|---|---|
| Without sound | 10 credits/s ($0.050/s) | Standard 720p resolution without an audio stream. |
| With sound | 13 credits/s ($0.065/s) | Includes native synchronized Foley sound effects, ambient audio, and multilingual dialogue. |
Best Use Cases
Creative Concept PrototypingTurn pitch outlines into dynamic 720p concept reels to evaluate pacing and lighting before production.
Social Feed ContentDeliver mobile-ready vertical short clips with native audio designed to maximize engagement and watch-through metrics.
Episodic PrevisualizationUse multi-shot scripting to convert screenplay beats into continuous visual animatics.
E-commerce Dynamic VisualsDescribe product demonstrations and ambient motion from text to generate eye-catching catalog hero videos.
Pro Tips
- Structure prompts hierarchically: character appearance first, followed by chronological actions, camera trajectory, and lighting atmosphere.
- When configuring multi-shot mode, verify that the sum of all shot durations precisely matches the top-level duration value.
- Mention explicit sound cues like footsteps splashing in rain or ringing chimes to help guide the audio diffusion pipeline.
- Specify smooth camera terms such as gentle pedestal up or slow lateral track for consistent mechanical stability.
- Allocate 2 to 4 seconds per shot in multi-shot sequences to give complex kinetic motions sufficient time to resolve physically.
Notes
- Single-shot mode requires a top-level prompt; activating multi-shot mode requires leaving the top-level prompt empty and defining each shot in the multi_prompt array.
- Multi-shot mode requires setting sound=true to guarantee cross-scene audiovisual cohesion.
- Generation runs asynchronously: tasks return a task_id immediately upon submission, with status accessible via polling or webhook notifications.
Kling O3 Standard Text to Video API FAQ
What does this endpoint generate?
Generate 3–15-second videos from text prompts with Kling O3 Standard. Supports single-shot and multi-shot generation. Explicitly set sound; multi-shot requires audio.
How do I set up multiple shots?
Required: explicitly send false for a single shot or true for multiple shots. Single-shot mode requires prompt. Multi-shot mode requires multi_prompt and sound=true, with no nonblank top-level prompt. Omitting this field is an error. Required when multi_shots=true; omit in single-shot mode. Supply at least one shot, each with a nonblank prompt of up to 2,500 characters and an integer duration of 1–12 seconds. Shot durations must sum to the top-level duration (3–15 seconds). Extra fields inside a shot are ignored.
How is the video aspect ratio determined?
Optional: 16:9 (landscape), 9:16 (portrait) or 1:1 (square). Sets the video aspect ratio.
How do I control audio?
Required: explicitly send true to generate audio or false for video without audio. Single-shot mode accepts either value; multi-shot mode requires true.
What are the prompt length limits?
The API validates prompts up to 2,500 characters. Some multi-shot generation requests have failed when a shot prompt exceeded 512 characters. We recommend keeping each shot prompt within 512 characters; this recommendation does not change the API validation limit.
How are duration and credits calculated?
Duration is 3–15 seconds. Video without audio costs 10 credits/s; video with audio costs 13 credits/s. Credits equal the top-level duration multiplied by the applicable rate.
What happens if generation fails or my request times out?
Invalid parameters are rejected before a generation task is created or credits are deducted. If an accepted generation task later reaches failed, its deducted credits are refunded. A client timeout alone does not mean the task failed; query its task_id before submitting again.