Photoreal indoor roller-derby footage on a teal track. One adult skater in a mustard jersey, black helmet and protective pads takes a tight left curve on quad skates. Knees bent, center of gravity low, wheels contacting the floor as weight shifts naturally. Low camera tracks smoothly alongside. Empty stands, overhead arena lights and realistic motion blur. Complete one clear turn without falling. No text, logos or watermarks.
Kling O3 Pro Text to Video API
kwaivgi/kling-video-o3-pro/text-to-videoGenerate 3–15-second videos from text prompts with Kling O3 Pro. Supports single-shot and multi-shot generation. Explicitly set sound; multi-shot requires audio.
Multi-shot requires sound. The form sets duration to the shot total.
Duration: 5 s (3-15)
Examples
REST API Reference
Quick Start
Submit a task and query its status.
Step 1: Set up authentication
Create an API key in the dashboard and attach Authorization: Bearer <API_KEY> when submitting a task.
- Submit Endpoint
- POST
https://api.vidgo.ai/api/generate/submit - Authorization Header
- Authorization: Bearer VIDGO_API_KEY
Step 2: Submit a task
POST /api/generate/submit: kwaivgi/kling-video-o3-pro/text-to-video
REQUEST_BODY=$(cat <<'JSON'
{
"model": "kwaivgi/kling-video-o3-pro/text-to-video",
"input": {
"duration": 5,
"sound": true,
"multi_shots": true,
"multi_prompt": [
{
"prompt": "Cinematic whimsical realism in a warm oak library. Medium two-shot: a small ivory robot with an oval face and amber eyes shelves a blue book, accidentally knocking one red book onto the floor. Beside it, an adult librarian wears a green cardigan and round glasses. Both remain visible under warm reading lamps. One distinct book thud against quiet room tone. Unmarked book covers. No text, logos or watermarks.",
"duration": 2
},
{
"prompt": "Cut closer to the same ivory robot and green-cardigan librarian in the same oak library. The librarian raises one finger to her lips. The robot tilts its head apologetically and softly says exactly \"Sorry.\" in a gentle robotic voice, synchronized with its small mouth light. Preserve their appearance, positions and warm lighting. End in an embarrassed pause; no music. No text, logos or watermarks.",
"duration": 3
}
],
"aspect_ratio": "16:9"
}
}
JSON
)
RESPONSE=$(curl --silent --show-error --fail-with-body \
--request POST \
--url "https://api.vidgo.ai/api/generate/submit" \
--header "Authorization: Bearer $VIDGO_API_KEY" \
--header "Content-Type: application/json" \
--data "$REQUEST_BODY")
CODE=$(printf '%s' "$RESPONSE" | jq -r '.code // empty')
if [ "$CODE" != "0" ] && [ "$CODE" != "200" ]; then
printf 'API error: %s
' "$RESPONSE" >&2
exit 1
fi
printf '%s
' "$RESPONSE"Step 3: Poll for completion
Poll with task_id while status is not_started or running, and stop at finished or failed. On success, read data.files[].file_url; on failure, read data.error_message.
Status Endpoint
GET https://api.vidgo.ai/api/generate/status/{task_id}Poll with task_id while status is not_started or running, and stop at finished or failed. On success, read data.files[].file_url; on failure, read data.error_message.
not_startedrunningfinishedfailed{
"code": 200,
"data": {
"task_id": "task-submitted-example",
"status": "not_started",
"created_time": "2026-09-22T00:00:00Z"
}
}{
"code": 200,
"data": {
"task_id": "BRDSCRKN4Q0WH4T6",
"status": "finished",
"files": [
{
"file_type": "video",
"file_url": "https://cdn.vidgo.ai/apis/models/kwaivgi/kling-video-o3-pro/text-to-video/v1/01/output.mp4"
}
],
"created_time": "2026-09-22T18:43:52",
"error_message": null,
"progress": 100
}
}Complete executable script
Expand to review an end-to-end script with automatic polling, error handling, and timeout safeguards.
set -euo pipefail
: "${VIDGO_API_KEY:?Set VIDGO_API_KEY in your environment}"
REQUEST_BODY=$(cat <<'JSON'
{
"model": "kwaivgi/kling-video-o3-pro/text-to-video",
"input": {
"duration": 5,
"sound": true,
"multi_shots": true,
"multi_prompt": [
{
"prompt": "Cinematic whimsical realism in a warm oak library. Medium two-shot: a small ivory robot with an oval face and amber eyes shelves a blue book, accidentally knocking one red book onto the floor. Beside it, an adult librarian wears a green cardigan and round glasses. Both remain visible under warm reading lamps. One distinct book thud against quiet room tone. Unmarked book covers. No text, logos or watermarks.",
"duration": 2
},
{
"prompt": "Cut closer to the same ivory robot and green-cardigan librarian in the same oak library. The librarian raises one finger to her lips. The robot tilts its head apologetically and softly says exactly \"Sorry.\" in a gentle robotic voice, synchronized with its small mouth light. Preserve their appearance, positions and warm lighting. End in an embarrassed pause; no music. No text, logos or watermarks.",
"duration": 3
}
],
"aspect_ratio": "16:9"
}
}
JSON
)
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
--request POST \
--url "https://api.vidgo.ai/api/generate/submit" \
--header "Authorization: Bearer $VIDGO_API_KEY" \
--header "Content-Type: application/json" \
--data "$REQUEST_BODY")
TASK_ID=$(printf '%s' "$SUBMIT_RESPONSE" | jq -r '.data.task_id // .task_id // empty')
BUSINESS_CODE=$(printf '%s' "$SUBMIT_RESPONSE" | jq -r '.code // empty')
if [ "$BUSINESS_CODE" != "0" ] && [ "$BUSINESS_CODE" != "200" ]; then
printf 'Submit failed:
%s
' "$SUBMIT_RESPONSE" >&2
exit 1
fi
if [ -z "$TASK_ID" ]; then
printf 'Submit response did not include task_id:
%s
' "$SUBMIT_RESPONSE" >&2
exit 1
fi
START_TIME=$(date +%s)
POLL_DELAY=2
while true; do
if [ $(( $(date +%s) - START_TIME )) -ge 600 ]; then
printf 'Timed out after 600 seconds
' >&2
exit 1
fi
STATUS_RESPONSE=$(curl --silent --show-error --fail-with-body \
--url "https://api.vidgo.ai/api/generate/status/$TASK_ID" \
--header "Authorization: Bearer $VIDGO_API_KEY")
STATUS=$(printf '%s' "$STATUS_RESPONSE" | jq -r '.data.status // .status // empty')
BUSINESS_CODE=$(printf '%s' "$STATUS_RESPONSE" | jq -r '.code // empty')
if [ "$BUSINESS_CODE" != "0" ] && [ "$BUSINESS_CODE" != "200" ]; then
printf 'Status request failed:
%s
' "$STATUS_RESPONSE" >&2
exit 1
fi
case "$STATUS" in
finished)
printf '%s' "$STATUS_RESPONSE" | jq -r '(.data.files // .files // [])[]?.file_url'
break
;;
failed)
printf '%s' "$STATUS_RESPONSE" | jq -r '.data.error_message // .error_message // "Generation failed"' >&2
exit 1
;;
not_started|running)
sleep "$POLL_DELAY"
if [ "$POLL_DELAY" -lt 10 ]; then POLL_DELAY=$((POLL_DELAY + 1)); fi
;;
*)
printf 'Unexpected task status: %s
' "$STATUS" >&2
exit 1
;;
esac
doneRequest Parameters (input object)
Put generation parameters inside input. Use standard JSON numbers and booleans. For compatibility, integer strings such as "5" are accepted. Boolean strings true/1/yes/y/on mean true; false/0/no/n/off mean false. These strings are case-insensitive and trimmed. Numeric 1 and 0 are also accepted for boolean fields. Numeric and boolean prompt values are converted to text; 0, false and null are treated as empty. Objects and arrays are not accepted as prompts. The API validates prompts up to 2,500 characters. Some multi-shot generation requests have failed when a shot prompt exceeded 512 characters. We recommend keeping each shot prompt within 512 characters; this recommendation does not change the API validation limit.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
| prompt | string | Conditional | - | Required when multi_shots=false. Describe the scene, action and camera movement. Use nonblank text, up to 2,500 characters after trimming. Omit prompt when multi_shots=true. |
| multi_shots | boolean | Yes | - | Required: explicitly send false for a single shot or true for multiple shots. Single-shot mode requires prompt. Multi-shot mode requires multi_prompt and sound=true, with no nonblank top-level prompt. Omitting this field is an error. |
| multi_prompt | array | Conditional | - | Required when multi_shots=true; omit in single-shot mode. Supply at least one shot, each with a nonblank prompt of up to 2,500 characters and an integer duration of 1–12 seconds. Shot durations must sum to the top-level duration (3–15 seconds). Extra fields inside a shot are ignored. |
| duration | integer | Yes | - | Required integer from 3 to 15 seconds. In multi-shot mode, this must equal the sum of all shot durations. Credits are calculated using this value. |
| sound | boolean | Yes | - | Required: explicitly send true to generate audio or false for video without audio. Single-shot mode accepts either value; multi-shot mode requires true. |
| aspect_ratio | string | No | - | Optional: 16:9 (landscape), 9:16 (portrait) or 1:1 (square). Sets the video aspect ratio. |
Response Fields (Status Query)
Details returned by GET /api/generate/status/{task_id}:
| Field | Type | Description |
|---|---|---|
| code | integer | HTTP/business response status code (200 indicates success). |
| data.task_id | string | Globally unique task identifier. |
| data.status | string | Task lifecycle state: not_started, running, finished, or failed. |
| data.files | array | Array of output assets containing file_url and file_type upon completion. |
| data.error_message | string | null | Error diagnostic details if the task status is failed. |
Task Lifecycle
Clients should poll status until reaching either the finished or failed terminal state:
not_startedQueued
runningGenerating
finishedReady
failedFailed
Polling & Error Handling
- Polling frequencyStart polling with a 2 to 3-second interval, gradually increasing to 5 seconds for extended takes.
- Network resiliencyTransient 5xx responses or timeouts do not signify task failure; retry status requests after a short backoff.
- Webhook callbacksProvide a top-level callback_url in your submission payload to receive completion notifications automatically.
Specifications
| Specification | Value | Description |
|---|---|---|
| Model | kwaivgi/kling-video-o3-pro/text-to-video | Kuaishou Kling O3 flagship 1080p text-to-video tier featuring Visual Chain-of-Thought reasoning and multi-shot sequence choreography. |
| Duration | 3-15 s | Required integer, 3–15 seconds. In multi-shot mode it must equal the sum of shot durations. Billing uses this duration. |
Kling O3 Pro Text to Video
Generate 3–15-second videos from text prompts with Kling O3 Pro. Supports single-shot and multi-shot generation. Explicitly set sound; multi-shot requires audio.
Why Choose This Model
1080p Broadcast QualityDelivers full high-definition clarity that captures skin pores, fine fabric weave, hair strands, and cinematic volumetric lighting.
Deep Visual Chain-of-ThoughtPre-plans scene geometry, character interactions, and dynamic camera choreography before rendering to prevent visual warping.
Advanced Multi-Shot SequencingDirect multiple camera setups in one generation while preserving character identity and color grading across cuts.
Studio-Grade Native AudioGenerates spatial ambient audio, synchronized contact sound effects, and realistic multilingual lip-sync in the same inference pass.
Cinematic Camera ChoreographyAccurately reproduces cranes, orbital tracks, dollies, and complex zooms for immersive, director-grade visual storytelling.
Parameters
| Parameter | Requirement | Description |
|---|---|---|
| prompt | Conditional | Required when multi_shots=false. Describe the scene, action and camera movement. Use nonblank text, up to 2,500 characters after trimming. Omit prompt when multi_shots=true. Default - |
| multi_shots | Yes | Required: explicitly send false for a single shot or true for multiple shots. Single-shot mode requires prompt. Multi-shot mode requires multi_prompt and sound=true, with no nonblank top-level prompt. Omitting this field is an error. Default - |
| multi_prompt | Conditional | Required when multi_shots=true; omit in single-shot mode. Supply at least one shot, each with a nonblank prompt of up to 2,500 characters and an integer duration of 1–12 seconds. Shot durations must sum to the top-level duration (3–15 seconds). Extra fields inside a shot are ignored. Default - |
| duration | Yes | Required integer from 3 to 15 seconds. In multi-shot mode, this must equal the sum of all shot durations. Credits are calculated using this value. Default - |
| sound | Yes | Required: explicitly send true to generate audio or false for video without audio. Single-shot mode accepts either value; multi-shot mode requires true. Default - |
| aspect_ratio | No | Optional: 16:9 (landscape), 9:16 (portrait) or 1:1 (square). Sets the video aspect ratio. Default - |
How to Use
Compose cinematic promptsDetail scene lighting, shot perspective, character wardrobe, and primary dramatic actions.
Structure camera and shotsChoose a continuous master take or activate multi-shot mode to sequence individual scene descriptions and durations.
Configure output parametersSelect an integer runtime between 3 and 15 seconds alongside 16:9, 9:16, or 1:1 framing.
Enable native audioSet sound to true to trigger the integrated audiovisual engine for dialogue and Foley effects.
Generate and downloadSubmit the task asynchronously, monitor progress using the task ID, and export clean 1080p MP4 footage.
Pricing
Credits = top-level duration × per-second rate. 1 credit = $0.005. If generation fails, consumed credits are automatically and fully refunded.
| Usage | Rate | Details |
|---|---|---|
| Without sound | 13 credits/s ($0.065/s) | Professional 1080p full high-definition tier without audio. |
| With sound | 16 credits/s ($0.080/s) | Full 1080p video paired with studio-grade native synchronized audio. |
Best Use Cases
Film and Series PrevisualizationTurn script passages into 1080p animatics to evaluate blocking, coverage, and narrative flow.
Commercial Brand CampaignsProduce polished, high-definition promotional videos ready for broadcast and high-resolution displays.
High-Value Short-Form DramaLeverage multi-shot scene transitions and lip-sync dialogue to generate compelling storytelling clips.
Animation and Game Cinematic ConceptsVisualize speculative lore, intricate creature kinetics, and dynamic sci-fi environments.
Pro Tips
- Incorporate professional filmmaking terms like "shot on 35mm lens", "subtle rim light", or "slow pedestal down" for precise aesthetic direction.
- In multi-character scenes, describe character positions and sequential actions distinctly to anchor spatial relationships.
- Keep the environmental tone consistent across multi-shot prompts (e.g., "moody neon rain reflections") to achieve natural, seamless editing.
- Wrap spoken dialogue lines in quotation marks within your prompt to guide accurate lip-sync phoneme generation.
- For complex physical action sequences, select at least 5 seconds of duration to give inertia and momentum sufficient frames to resolve naturally.
Notes
- The Pro tier is tailored for 1080p rendering and utilizes deeper reasoning iterations than the Standard tier.
- Enabling multi_shots requires sound=true and mandates that the top-level prompt be omitted.
- API operations are asynchronous: query task status using the returned task_id until a terminal state is reached.
Kling O3 Pro Text to Video API FAQ
What does this endpoint generate?
Generate 3–15-second videos from text prompts with Kling O3 Pro. Supports single-shot and multi-shot generation. Explicitly set sound; multi-shot requires audio.
How do I set up multiple shots?
Required: explicitly send false for a single shot or true for multiple shots. Single-shot mode requires prompt. Multi-shot mode requires multi_prompt and sound=true, with no nonblank top-level prompt. Omitting this field is an error. Required when multi_shots=true; omit in single-shot mode. Supply at least one shot, each with a nonblank prompt of up to 2,500 characters and an integer duration of 1–12 seconds. Shot durations must sum to the top-level duration (3–15 seconds). Extra fields inside a shot are ignored.
How is the video aspect ratio determined?
Optional: 16:9 (landscape), 9:16 (portrait) or 1:1 (square). Sets the video aspect ratio.
How do I control audio?
Required: explicitly send true to generate audio or false for video without audio. Single-shot mode accepts either value; multi-shot mode requires true.
What are the prompt length limits?
The API validates prompts up to 2,500 characters. Some multi-shot generation requests have failed when a shot prompt exceeded 512 characters. We recommend keeping each shot prompt within 512 characters; this recommendation does not change the API validation limit.
How are duration and credits calculated?
Duration is 3–15 seconds. Video without audio costs 13 credits/s; video with audio costs 16 credits/s. Credits equal the top-level duration multiplied by the applicable rate.
What happens if generation fails or my request times out?
Invalid parameters are rejected before a generation task is created or credits are deducted. If an accepted generation task later reaches failed, its deducted credits are refunded. A client timeout alone does not mean the task failed; query its task_id before submitting again.