Kling 3.0 Turbo Pro Text to Video API

kwaivgi/kling-v3-turbo-pro/text-to-video

Kling 3.0 Turbo Pro Text to Video turns text prompts into 3–15 second flagship 1080p video with refined textures and physics dynamics, native audio-visual lip sync, and multi-shot scripts of up to 6 shots. It keeps character appearance and scene continuity across cuts while aligning dialogue lip sync and ambient sound with on-screen action. With multi_prompt, omitting top-level duration also allows shot totals of 1–2 seconds.

Input
0/2500
Output
Idle

Your generated video will appear here

Add your prompt and required media, review the settings, then click Run.

Kling 3.0 Turbo Pro · 5 sec × 22/sec = $0.550 (110 credits)
Continue with

Examples

Low drone shot skimming over golden sunset waves crashing against dark coastal rocks, spray catching the warm light, cinematic color grading, epic seascape

A great grey owl in a snowy forest turning its head to look directly at the camera, snowflakes drifting down, soft winter light, wildlife documentary close-up

A green sea turtle gliding over a vibrant coral reef, sun rays piercing through crystal clear blue water, schools of small fish scattering, underwater documentary footage

Kling 3.0 Turbo Pro Text to Video

Kling 3.0 Turbo Pro Text to Video is built for delivery-ready shorts: from natural-language prompts alone it produces 3–15 second flagship 1080p motion with synchronized multilingual dialogue lip sync and ambient sound. With finer texture rendering and physics dynamics, plus multi_prompt storyboards of up to 6 shots, it fits brand ads, vertical social narratives, and performance clips that need tight mouth alignment. With multi_prompt, omitting top-level duration also allows shot totals of 1–2 seconds.

Why Choose This?

  • Flagship 1080p delivery qualityThe Pro tier locks studio-grade full HD 1080p with clearer edges and material detail, ready for editorial finishing and final delivery.

  • Refined textures and physics dynamicsStrengthens cloth, rigid-body, and environmental interaction so subject motion and lighting stay natural as the camera advances.

  • Native audio-visual lip syncSynthesizes matching audio with the picture and aligns dialogue mouth shapes to speech—ideal for talking-heads, dialogue, and reaction shots.

  • Multi-shot scripts up to 6 shotsUse multi_prompt to define prompts and durations for 1–6 shots in one request; the model sequences cuts along the timeline.

  • Cross-shot subject and scene continuityKeeps faces, wardrobe, and setting cues consistent through framing changes and camera moves for complete short-form arcs.

  • Flexible 3–15s duration and framingInteger durations default to 5 seconds, with 16:9, 9:16, and 1:1 aspect ratios for landscape finals and vertical social placements.

Parameters

ParameterRequirementDescription
promptOptional

String, 1–2500 characters. Mutually exclusive with multi_prompt.

durationOptional

An explicit duration must be 3–15 seconds and equal the shot total when multi_prompt is set. Omit it to use the shot total (1–15 seconds); without multi_prompt it defaults to 5 seconds.

DefaultShot total / 5
aspect_ratioOptional

String. Text-to-video framing; the Playground preselects 16:9.

Default16:99:161:1
multi_promptOptional

1–6 shots, each with a nonblank prompt and optional duration of 1–15 seconds (default 5). Total duration must not exceed 15 seconds. Mutually exclusive with a nonblank prompt.

How to Use

  1. Write the scene promptIn prompt, specify subject look, action, camera moves, and the sound cues you want. For multi-shot work, switch to multi_prompt and describe each shot separately.

  2. Set output durationWithout multi_prompt, duration is an integer from 3 to 15 seconds, defaulting to 5. With multi_prompt, omit top-level duration to use the shot total (1–15 seconds); if supplied, it must be 3–15 seconds and equal that total.

  3. Choose aspect ratioPick 16:9, 9:16, or 1:1 for your delivery layout; the Playground defaults to 16:9. For vertical social content, select 9:16 directly.

  4. Confirm 1080p deliveryThe Pro tier always outputs 1080p—no separate resolution field is required—so results are ready for finals and HD review.

  5. Storyboard multi-shot (optional)For multi-shot narratives, enable multi_prompt with 1–6 shots, keep it mutually exclusive with prompt, and keep total duration at or under 15 seconds.

  6. Review the cost and runCheck the cost shown on the Run button (output seconds × 22 credits/sec), finish your settings, then click Run.

  7. Preview and download the videoWhen the task finishes, preview picture and synchronized audio in the output panel, then select Download video to save the result.

Pricing

Billed by output video seconds. 1 credit = $0.005. The Pro tier is fixed at 1080p.

UsageRateDetails
1080p22 credits/sec ($0.11/sec)Default 5s at 1080p is 110 credits ($0.55). Official comparison $0.1375/sec.

Best Use Cases

  • Brand ads and product finalsDescribe product look, camera path, and pacing in detail to generate 1080p promo clips ready for editorial.

  • Multi-shot advertising scriptsUse multi_prompt to stage close-ups, tracking beats, and closing frames in one complete ad arc.

  • Vertical social storytellingChoose 9:16 with plot and sound cues to produce high-definition vertical clips with lip-synced dialogue.

  • Talking-head and performance shotsLean on native audio-visual lip sync for natural talking-heads, dialogue, and reaction performances.

  • E-commerce product motionWrite material reflections, hand interaction, and environmental motion to create credible product dynamics.

Pro Tips

  • Reuse the same appearance and wardrobe wording for a character across multi-shot scripts to stabilize identity between cuts.
  • State framing, camera move, and action beat per shot—for example, “3s extreme close-up of the sole landing,” then “low-angle side follow as the runner accelerates.”
  • For lip sync, name the spoken language, emotion, and speech rhythm in the prompt so audio and mouth shapes stay aligned.
  • Pick 9:16 for vertical feeds; prefer 16:9 for landscape ads and widescreen storytelling.
  • Use multi_prompt for 1–6 shots, each with a required prompt and optional duration defaulting to 5 seconds. The total must not exceed 15 seconds. Top-level duration can be omitted; if supplied, it must equal the total and be 3–15 seconds.

Usage notes

  • Kling 3.0 Turbo Pro Text to Video is driven by a text prompt with optional duration and aspect_ratio; use multi_prompt for storyboards—the two inputs are mutually exclusive.
  • The Pro tier is fixed at 1080p and billed at 22 credits per output second; the default 5-second clip costs 110 credits ($0.55).
  • After an API submission, save the returned task_id to query progress and retrieve the final media URL.
  • Generated clips are delivered as standard video files ready for playback and editing software.

Kling 3.0 Turbo Pro Text to Video API frequently asked questions

What is the Kling 3.0 Turbo Pro Text to Video API?

Kling 3.0 Turbo Pro Text to Video is a Kuaishou (Kling AI) flagship Turbo Pro model for generating video from text prompts. It creates flagship 1080p clips with refined textures and physics dynamics, native audio-visual lip sync, and multi_prompt storyboards of up to 6 shots. Built on the Kling 3.0 family's multi-shot storytelling and joint audiovisual generation, it follows camera and sound design in the prompt while keeping character appearance, scene continuity, and dialogue lip sync aligned. You can call it programmatically or try it from the playground above. Without multi_prompt, duration is 3–15 seconds. With multi_prompt, omitting top-level duration also allows a shot total of 1 or 2 seconds.

Is Kling 3.0 Turbo Pro Text to Video 1080p quality delivery-ready?

Yes. The Pro tier always outputs studio-grade full HD 1080p with finer texture and edge detail, so clips can move straight into editorial as final assets. Resolution is fixed by the tier—no separate resolution field is required. Full specs are listed in the Parameters and Endpoint limits sections on this page.

How does Kling 3.0 Turbo Pro Text to Video use multi_prompt storyboards?

multi_prompt accepts 1–6 shots. Each shot needs a nonblank prompt and an optional integer duration of 1–15 seconds, defaulting to 5. The sum must not exceed 15 seconds. Omit top-level duration to use that sum, including totals of 1 or 2 seconds. If supplied, top-level duration must be an integer from 3 to 15 and equal the sum. multi_prompt cannot be combined with a nonblank top-level prompt. Edit storyboards in JSON mode; switching to the form keeps JSON mode and all shot settings.

Does Kling 3.0 Turbo Pro Text to Video support native lip sync?

Yes. Results include synchronized audio with the picture and mouth-shape alignment for multilingual dialogue—well suited to talking-heads, conversation, and reaction shots. Name the spoken language, emotion, and speech rhythm in the prompt to further guide the soundtrack and lip performance.

Which aspect ratio fits Kling 3.0 Turbo Pro Text to Video?

Choose among 16:9, 9:16, and 1:1; the Playground defaults to 16:9. Prefer 16:9 for landscape ads and finals, select 9:16 for TikTok / Reels / Shorts, and use 1:1 for square covers or social posts. Full options are listed in the Parameters section on this page.

When should I pick Kling 3.0 Turbo Pro Text to Video over the Base model?

Choose this Pro text-to-video tier when you need flagship 1080p delivery, finer textures, and lip-synced shorts. For faster drafts, higher throughput, and a lower per-second cost, switch to the same-family Standard 720p text-to-video tier. Full Kling 3.0 Base emphasizes higher-fidelity cinematic hero shots—compare family options under Related Models on this page.

How is Kling 3.0 Turbo Pro Text to Video billed per second?

Billing is by output video seconds at 22 credits/sec ($0.11/sec, 1 credit = $0.005). The default 5-second clip costs 110 credits ($0.55); longer durations cost more. See the Pricing section on this page for the full rate.