Kling O3 Pro Text to Video API

kwaivgi/kling-video-o3-pro/text-to-video

Generate 3–15-second videos from text prompts with Kling O3 Pro. Supports single-shot and multi-shot generation. Explicitly set sound; multi-shot requires audio.

Input

Multi-shot requires sound. The form sets duration to the shot total.

410/2,500
399/2,500

Duration: 5 s (3-15)

Output
Ready
5 s × 16 credits/s = 80 credits ($0.400)
Continue using

Examples

Photoreal indoor roller-derby footage on a teal track. One adult skater in a mustard jersey, black helmet and protective pads takes a tight left curve on quad skates. Knees bent, center of gravity low, wheels contacting the floor as weight shifts naturally. Low camera tracks smoothly alongside. Empty stands, overhead arena lights and realistic motion blur. Complete one clear turn without falling. No text, logos or watermarks.

Photoreal aerial shot gliding alongside a vast cumulonimbus tower above a cloud sea at sunrise. Cool sculpted shadows contrast with an amber horizon. One internal lightning pulse briefly illuminates branching shapes deep in the cloud, then fades back to dawn light. Volumetric depth, delicate wisps and coherent lighting. No visible aircraft, ground, people or buildings. No text, logos or watermarks.

Kling O3 Pro Text to Video

Generate 3–15-second videos from text prompts with Kling O3 Pro. Supports single-shot and multi-shot generation. Explicitly set sound; multi-shot requires audio.

Why Choose This Model

  • 1080p Broadcast QualityDelivers full high-definition clarity that captures skin pores, fine fabric weave, hair strands, and cinematic volumetric lighting.

  • Deep Visual Chain-of-ThoughtPre-plans scene geometry, character interactions, and dynamic camera choreography before rendering to prevent visual warping.

  • Advanced Multi-Shot SequencingDirect multiple camera setups in one generation while preserving character identity and color grading across cuts.

  • Studio-Grade Native AudioGenerates spatial ambient audio, synchronized contact sound effects, and realistic multilingual lip-sync in the same inference pass.

  • Cinematic Camera ChoreographyAccurately reproduces cranes, orbital tracks, dollies, and complex zooms for immersive, director-grade visual storytelling.

Parameters

ParameterRequirementDescription
promptConditional

Required when multi_shots=false. Describe the scene, action and camera movement. Use nonblank text, up to 2,500 characters after trimming. Omit prompt when multi_shots=true.

Default-
multi_shotsYes

Required: explicitly send false for a single shot or true for multiple shots. Single-shot mode requires prompt. Multi-shot mode requires multi_prompt and sound=true, with no nonblank top-level prompt. Omitting this field is an error.

Default-
multi_promptConditional

Required when multi_shots=true; omit in single-shot mode. Supply at least one shot, each with a nonblank prompt of up to 2,500 characters and an integer duration of 1–12 seconds. Shot durations must sum to the top-level duration (3–15 seconds). Extra fields inside a shot are ignored.

Default-
durationYes

Required integer from 3 to 15 seconds. In multi-shot mode, this must equal the sum of all shot durations. Credits are calculated using this value.

Default-
soundYes

Required: explicitly send true to generate audio or false for video without audio. Single-shot mode accepts either value; multi-shot mode requires true.

Default-
aspect_ratioNo

Optional: 16:9 (landscape), 9:16 (portrait) or 1:1 (square). Sets the video aspect ratio.

Default-

How to Use

  1. Compose cinematic promptsDetail scene lighting, shot perspective, character wardrobe, and primary dramatic actions.

  2. Structure camera and shotsChoose a continuous master take or activate multi-shot mode to sequence individual scene descriptions and durations.

  3. Configure output parametersSelect an integer runtime between 3 and 15 seconds alongside 16:9, 9:16, or 1:1 framing.

  4. Enable native audioSet sound to true to trigger the integrated audiovisual engine for dialogue and Foley effects.

  5. Generate and downloadSubmit the task asynchronously, monitor progress using the task ID, and export clean 1080p MP4 footage.

Pricing

Credits = top-level duration × per-second rate. 1 credit = $0.005. If generation fails, consumed credits are automatically and fully refunded.

UsageRateDetails
Without sound13 credits/s ($0.065/s)Professional 1080p full high-definition tier without audio.
With sound16 credits/s ($0.080/s)Full 1080p video paired with studio-grade native synchronized audio.

Best Use Cases

  • Film and Series PrevisualizationTurn script passages into 1080p animatics to evaluate blocking, coverage, and narrative flow.

  • Commercial Brand CampaignsProduce polished, high-definition promotional videos ready for broadcast and high-resolution displays.

  • High-Value Short-Form DramaLeverage multi-shot scene transitions and lip-sync dialogue to generate compelling storytelling clips.

  • Animation and Game Cinematic ConceptsVisualize speculative lore, intricate creature kinetics, and dynamic sci-fi environments.

Pro Tips

  • Incorporate professional filmmaking terms like "shot on 35mm lens", "subtle rim light", or "slow pedestal down" for precise aesthetic direction.
  • In multi-character scenes, describe character positions and sequential actions distinctly to anchor spatial relationships.
  • Keep the environmental tone consistent across multi-shot prompts (e.g., "moody neon rain reflections") to achieve natural, seamless editing.
  • Wrap spoken dialogue lines in quotation marks within your prompt to guide accurate lip-sync phoneme generation.
  • For complex physical action sequences, select at least 5 seconds of duration to give inertia and momentum sufficient frames to resolve naturally.

Notes

  • The Pro tier is tailored for 1080p rendering and utilizes deeper reasoning iterations than the Standard tier.
  • Enabling multi_shots requires sound=true and mandates that the top-level prompt be omitted.
  • API operations are asynchronous: query task status using the returned task_id until a terminal state is reached.

Kling O3 Pro Text to Video API FAQ

What does this endpoint generate?

Generate 3–15-second videos from text prompts with Kling O3 Pro. Supports single-shot and multi-shot generation. Explicitly set sound; multi-shot requires audio.

How do I set up multiple shots?

Required: explicitly send false for a single shot or true for multiple shots. Single-shot mode requires prompt. Multi-shot mode requires multi_prompt and sound=true, with no nonblank top-level prompt. Omitting this field is an error. Required when multi_shots=true; omit in single-shot mode. Supply at least one shot, each with a nonblank prompt of up to 2,500 characters and an integer duration of 1–12 seconds. Shot durations must sum to the top-level duration (3–15 seconds). Extra fields inside a shot are ignored.

How is the video aspect ratio determined?

Optional: 16:9 (landscape), 9:16 (portrait) or 1:1 (square). Sets the video aspect ratio.

How do I control audio?

Required: explicitly send true to generate audio or false for video without audio. Single-shot mode accepts either value; multi-shot mode requires true.

What are the prompt length limits?

The API validates prompts up to 2,500 characters. Some multi-shot generation requests have failed when a shot prompt exceeded 512 characters. We recommend keeping each shot prompt within 512 characters; this recommendation does not change the API validation limit.

How are duration and credits calculated?

Duration is 3–15 seconds. Video without audio costs 13 credits/s; video with audio costs 16 credits/s. Credits equal the top-level duration multiplied by the applicable rate.

What happens if generation fails or my request times out?

Invalid parameters are rejected before a generation task is created or credits are deducted. If an accepted generation task later reaches failed, its deducted credits are refunded. A client timeout alone does not mean the task failed; query its task_id before submitting again.