One continuous three-second medium close shot in a modest basement rehearsal room. An adult drummer makes one clear controlled strike on the snare drum, the stick rebounds naturally and his shoulders settle. Keep both hands anatomically consistent and the drum fixed. Warm practical lighting, documentary realism, slight handheld breathing without a cut. Audio: one crisp synchronized snare hit and a short natural room decay, quiet room tone, no music, no speech. No lettering, brands or watermark.
Kling 3.0 Standard Text to Video API
kwaivgi/kling-v3.0-std/text-to-videoKling 3.0 Standard Text to Video creates 720p clips from prompts, with timed shots and native audio. Direct camera movement and describe recurring characters to carry a scene through each cut.
457/2,500
Examples
REST API
Quick Start
Submit a Kling 3.0 Standard Text to Video request and use its task identifier to retrieve the video.
Authentication
Send your Vidgo API key in the Authorization header as a Bearer token.
- Submit endpoint
- POST
https://api.vidgo.ai/api/generate/submit - Authentication
- Authorization: Bearer YOUR_API_KEY
Submit Request
Send the model identifier and scene prompt inside the JSON request. Change the prompt to describe your scene.
REQUEST_BODY=$(cat <<'JSON'
{
"model": "kwaivgi/kling-v3.0-std/text-to-video",
"input": {
"prompt": "A ceramic cup sits beside a window in morning light. The camera slowly moves closer as steam curls above the rim. A quiet room tone accompanies the scene.",
"duration": 5,
"multi_shots": false,
"sound": true,
"aspect_ratio": "16:9"
}
}
JSON
)
RESPONSE=$(curl --silent --show-error --fail-with-body \
--request POST \
--url "https://api.vidgo.ai/api/generate/submit" \
--header "Authorization: Bearer $VIDGO_API_KEY" \
--header "Content-Type: application/json" \
--data "$REQUEST_BODY")
CODE=$(printf '%s' "$RESPONSE" | jq -r '.code // empty')
if [ "$CODE" != "0" ] && [ "$CODE" != "200" ]; then
printf 'API error: %s
' "$RESPONSE" >&2
exit 1
fi
printf '%s
' "$RESPONSE"Query Task Status
Replace {task_id} in the status URL with the identifier returned on submission.
Status endpoint
https://api.vidgo.ai/api/generate/status/{task_id}Query this URL with the returned task_id. Continue until finished or failed; a completed video is listed in data.files.
not_startedrunningfinishedfailed{
"code": 200,
"data": {
"task_id": "example-task-id",
"status": "not_started",
"created_time": "2026-09-27T00:00:00Z"
}
}{
"code": 200,
"data": {
"task_id": "example-task-id",
"status": "finished",
"created_time": "2026-09-27T00:00:00Z",
"files": [
{
"file_type": "video",
"file_url": "https://example.com/generated-video.mp4"
}
]
}
}Complete Example
When status is finished, read file_url from video entries in files. Task IDs and output URLs below illustrate the response format.
set -euo pipefail
: "${VIDGO_API_KEY:?Set VIDGO_API_KEY in your environment}"
REQUEST_BODY=$(cat <<'JSON'
{
"model": "kwaivgi/kling-v3.0-std/text-to-video",
"input": {
"prompt": "A ceramic cup sits beside a window in morning light. The camera slowly moves closer as steam curls above the rim. A quiet room tone accompanies the scene.",
"duration": 5,
"multi_shots": false,
"sound": true,
"aspect_ratio": "16:9"
}
}
JSON
)
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
--request POST \
--url "https://api.vidgo.ai/api/generate/submit" \
--header "Authorization: Bearer $VIDGO_API_KEY" \
--header "Content-Type: application/json" \
--data "$REQUEST_BODY")
TASK_ID=$(printf '%s' "$SUBMIT_RESPONSE" | jq -r '.data.task_id // .task_id // empty')
BUSINESS_CODE=$(printf '%s' "$SUBMIT_RESPONSE" | jq -r '.code // empty')
if [ "$BUSINESS_CODE" != "0" ] && [ "$BUSINESS_CODE" != "200" ]; then
printf 'Submit failed:
%s
' "$SUBMIT_RESPONSE" >&2
exit 1
fi
if [ -z "$TASK_ID" ]; then
printf 'Submit response did not include task_id:
%s
' "$SUBMIT_RESPONSE" >&2
exit 1
fi
START_TIME=$(date +%s)
POLL_DELAY=2
while true; do
if [ $(( $(date +%s) - START_TIME )) -ge 600 ]; then
printf 'Timed out after 600 seconds
' >&2
exit 1
fi
STATUS_RESPONSE=$(curl --silent --show-error --fail-with-body \
--url "https://api.vidgo.ai/api/generate/status/$TASK_ID" \
--header "Authorization: Bearer $VIDGO_API_KEY")
STATUS=$(printf '%s' "$STATUS_RESPONSE" | jq -r '.data.status // .status // empty')
BUSINESS_CODE=$(printf '%s' "$STATUS_RESPONSE" | jq -r '.code // empty')
if [ "$BUSINESS_CODE" != "0" ] && [ "$BUSINESS_CODE" != "200" ]; then
printf 'Status request failed:
%s
' "$STATUS_RESPONSE" >&2
exit 1
fi
case "$STATUS" in
finished)
printf '%s' "$STATUS_RESPONSE" | jq -r '(.data.files // .files // [])[]?.file_url'
break
;;
failed)
printf '%s' "$STATUS_RESPONSE" | jq -r '.data.error_message // .error_message // "Generation failed"' >&2
exit 1
;;
not_started|running)
sleep "$POLL_DELAY"
if [ "$POLL_DELAY" -lt 10 ]; then POLL_DELAY=$((POLL_DELAY + 1)); fi
;;
*)
printf 'Unexpected task status: %s
' "$STATUS" >&2
exit 1
;;
esac
doneRequest Parameters
Choose a single prompt or a sequence of shots, then set duration and sound for this endpoint.
| Field | Type | Requirement | Default | Usage |
|---|---|---|---|---|
| prompt | string | When multi_shots=false | Set explicitly | Describe the subject, scene and camera direction in 1–2,500 characters. For multiple shots, write each scene in multi_prompt. |
| multi_shots | boolean | Optional | false | Choose false for one prompt, or true for separately timed shots with multi_prompt and sound=true. |
| multi_prompt | object[] | When multi_shots=true | Set explicitly | Add at least one object with prompt and duration. Each prompt uses 1–2,500 characters; each shot lasts 1–12 integer seconds. Shot times add up to duration. |
| duration | integer | Required | Set explicitly | Set the total clip length to an integer from 3 to 15 seconds. For multiple shots, use the sum of their durations. |
| sound | boolean | Optional | true | Use true for generated audio. A silent single shot uses false. Multiple shots use true. |
| aspect_ratio | string | Optional | 1:1 | Choose 1:1, 16:9 or 9:16 to frame square, landscape or portrait scenes. |
Response Fields
Submission returns a task identifier. Query its status to retrieve generated video URLs.
| Field | Type | Usage |
|---|---|---|
| code | integer | Response code: 0 or 200 indicates a successful API operation. |
| data.task_id | string | Identifier returned on submission; use it to query the same task. |
| data.status | string | One of not_started, running, finished or failed. A successful submission is followed by status checks. |
| data.created_time | string | Task creation time returned by the service. |
| data.files | array | Generated media entries returned when the task finishes. |
| data.files[].file_type | string | Use entries with file_type=video for the generated clip. |
| data.files[].file_url | string | Video URL for playback or download. |
| data.error_message | string | null | Read the task error detail when status=failed. |
Task Lifecycle
Keep the returned task_id and follow status until finished or failed.
not_startedThe request is accepted and waiting to start. Keep its task_id for the next status check.
runningThe video is being generated. Continue checking the same task.
finishedGeneration is complete. Read the video entries in files.
failedGeneration ended with an error. Review error_message before submitting an adjusted request.
Error Handling
- Check timing before submittingFor multiple shots, use sound=true and ensure individual durations sum to a total of 3–15 seconds.
- Keep reference URLs accessibleUse public HTTP(S) media URLs that remain reachable while the task is processed.
- Resume interrupted status checksIf a status request fails, retry the status check using the same task_id. Submit a new generation only when you intend to create another clip.
- Read the returned errorFor a failed task, inspect error_message and adjust the relevant input before trying again.
Specifications
| Specification | Value | Usage |
|---|---|---|
| Input mode | Text to Video | One scene prompt or a sequence of individually timed shot prompts. |
| Resolution | 720p | 720p video output for this tier. |
| Clip duration | 3–15 seconds | Whole seconds; each shot in a sequence uses 1–12 seconds. |
| Audio | sound | Generate audio with true; single-shot silent generation uses false. |
| Model identifier | kwaivgi/kling-v3.0-std/text-to-video | Public model value for this request. |
Related Models
Kling 3.0 Standard Text to Video API frequently asked questions
What is the Kling 3.0 Standard Text to Video API?
Kling 3.0 Standard Text to Video is a Kuaishou model for generating scenes from written prompts. It creates 720p footage with audio and individually timed shots. Joint audio and video generation connects sound with the scene, while per-shot prompts guide its action and continuity. You can call it programmatically or try it from the playground above.
How are multi-shot scenes timed in Kling 3.0 Standard Text to Video?
Enable multi_shots and sound, then supply a prompt and an integer duration of 1–12 seconds for each shot. Set duration to the sum of the shot times, between 3 and 15 seconds.
Can Kling 3.0 Standard Text to Video generate silent storyboards?
Use sound=false with multi_shots=false to generate a silent single-shot draft. To plan a sequence in one generation, enable multi_shots and sound and supply individual shot prompts.
How do camera prompts guide Kling 3.0 Standard Text to Video?
State the shot size, camera direction and subject action together, such as a close-up followed by a slow pullback. In a sequence, assign each camera move to its own shot prompt and duration.
How do I guide scene continuity in Kling 3.0 Standard Text to Video?
Repeat the same character traits, clothing, setting and lighting across shot prompts. Change the action or camera angle deliberately so the next shot has a clear relationship to the previous one.
How long can a Kling 3.0 Standard Text to Video clip be?
Set duration to any integer from 3 to 15 seconds. For a sequence, divide that time among shot prompts; each individual shot uses 1–12 seconds.
When should I choose Kling 3.0 Standard over Kling 3.0 Pro?
Choose Standard for 720p scene exploration at a lower per-second rate. Choose Pro when the output needs 1080p detail; both offer audio and individually timed shots.
How is a 15-second Kling 3.0 Standard sequence billed?
A multi-shot sequence uses sound=true at $0.195 per second, so 15 seconds costs $2.925. Divide the total among shots of 1–12 seconds; their durations must sum to 15.















