Kling 3.0 Turbo Standard Text to Video API

kwaivgi/kling-v3-turbo-std/text-to-video

Kling 3.0 Turbo Standard Text to Video turns text prompts into 3–15 second 720p video with high-throughput generation, native lip-synced multilingual audio, and multi_prompt storyboarding for 1–6 shots. It keeps character appearance and scene continuity across cuts while aligning motion with dialogue and ambient sound. With multi_prompt, omitting top-level duration also allows shot totals of 1–2 seconds.

Input
0/2500
Output
Idle

Your generated video will appear here

Add your prompt and required media, review the settings, then click Run.

Kling 3.0 Turbo Standard · 5 sec × 17/sec = $0.425 (85 credits)
Continue with

Examples

Macro slow-motion shot of a raindrop falling onto a mossy forest floor, ripples expanding in a tiny puddle, spore dust drifting in a shaft of soft light, ultra detailed, nature documentary style

Rainy night in a neon-lit city street, a pedestrian with a transparent umbrella crosses the wet zebra crossing, reflections shimmering on the asphalt, cinematic vertical shot

A chef tossing a wok over intense heat, flames leaping up and wrapping around the ingredients, sparks and smoke, dynamic close-up in a professional kitchen

Kling 3.0 Turbo Standard Text to Video

Kling 3.0 Turbo Standard Text to Video is Kuaishou Kling AI’s high-throughput model for agile video creation from text. From natural-language prompts alone, it outputs 3–15 second 720p clips with multilingual lip-synced dialogue and ambient audio. With native multi_prompt storyboarding, a single request can sequence up to 6 coherent shots—ideal for social short-form batches, rapid concept drafts, and digital marketing content. With multi_prompt, omitting top-level duration also allows shot totals of 1–2 seconds.

Why Choose This?

  • High-throughput, low-latency iterationTuned for volume production so you can validate many prompt variants quickly with shorter wait times per task.

  • Native audio with faithful lip syncSynthesizes matching audio with the picture, aligning dialogue lip sync and natural speech so you skip separate dubbing and A/V alignment.

  • Multi-shot storyboarding up to 6 shotsUse multi_prompt to define 1–6 shots with per-shot prompts and durations; the model schedules pacing and cuts in one run.

  • Strong subject continuity across cutsMaintains face, wardrobe, and setting consistency as framing and camera moves change between shots.

  • Flexible 3–15 second integer durationPick any integer length from 3 to 15 seconds—tight 3-second hooks or fuller short narratives—with a default of 5 seconds.

  • Cost-efficient fixed 720p Standard rateDelivers clear 720p at 17 credits per second ($0.085/sec), lowering the cost of large-scale content production.

Parameters

ParameterRequirementDescription
promptOptional

String, 1–2500 characters. Mutually exclusive with multi_prompt.

durationOptional

An explicit duration must be 3–15 seconds and equal the shot total when multi_prompt is set. Omit it to use the shot total (1–15 seconds); without multi_prompt it defaults to 5 seconds.

DefaultShot total / 5
aspect_ratioOptional

String. Text-to-video framing; the Playground preselects 16:9.

Default16:99:161:1
multi_promptOptional

1–6 shots, each with a nonblank prompt and optional duration of 1–15 seconds (default 5). Total duration must not exceed 15 seconds. Mutually exclusive with a nonblank prompt.

How to Use

  1. Write the scene promptIn prompt, describe subject action, camera, lighting, and dialogue or ambience cues. For a single narrative block, use prompt only—do not submit it together with multi_prompt.

  2. Configure multi_prompt storyboardingFor multi-shot runs, switch to multi_prompt and add 1–6 shot objects, each with its own prompt and optional duration.

  3. Set output durationWithout multi_prompt, duration is an integer from 3 to 15 seconds, defaulting to 5. With multi_prompt, omit top-level duration to use the shot total (1–15 seconds); if supplied, it must be 3–15 seconds and equal that total.

  4. Select aspect ratioPick 16:9, 9:16, or 1:1 to match distribution layout; the Playground preselects 16:9.

  5. Verify total shot lengthUse multi_prompt for 1–6 shots, each with a required prompt and optional duration defaulting to 5 seconds. The total must not exceed 15 seconds. Top-level duration can be omitted; if supplied, it must equal the total and be 3–15 seconds.

  6. Review the cost and runCheck the cost shown on the Run button, finish the prompt and settings, then click Run.

  7. Preview and download the videoWhen the task finishes, preview picture and synchronized audio in the output panel, then select Download video to save the result.

Pricing

Billed by output video seconds. 1 credit = $0.005. The Standard tier is fixed at 720p.

UsageRateDetails
720p17 credits/sec ($0.085/sec)Default 5s at 720p is 85 credits ($0.425). Official comparison $0.106/sec.

Best Use Cases

  • High-volume social short-formBatch 9:16 clips at short durations for TikTok, Reels, and Shorts workflows that need fast turnaround.

  • Multi-shot ad sequences in one runStoryboard openers, tracking beats, and closers with multi_prompt while keeping product and talent consistent across cuts.

  • Rapid concept prototypingIterate prompts and pacing at 720p Standard before promoting selected ideas to higher-resolution delivery tiers.

  • Lip-synced talking performancesDirect talking-heads, dialogue, or reaction shots with speech cues so mouth motion tracks multilingual delivery.

  • E-commerce and brand promosDescribe product benefits and scene motion in text to produce marketing clips with ambient sound for horizontal or vertical placement.

Pro Tips

  • Specify framing, camera move, and lighting—e.g., “low-angle tracking shot with side light”—so motion and space land clearly.
  • With multi_prompt, reuse the same character appearance and wardrobe wording across shots to lock continuity.
  • Use multi_prompt for 1–6 shots, each with a required prompt and optional duration defaulting to 5 seconds. The total must not exceed 15 seconds. Top-level duration can be omitted; if supplied, it must equal the total and be 3–15 seconds.
  • For vertical social, choose 9:16 and state where the subject sits in the tall frame.
  • When you need lip sync, add language and tone cues in the prompt to guide native audiovisual alignment.

Usage notes

  • Kling 3.0 Turbo Standard Text to Video is driven by prompt or multi_prompt (mutually exclusive), with configurable duration and aspect_ratio.
  • The Standard tier is fixed at 720p and includes native audio with lip sync in the result—no separate sound toggle field.
  • After an API submission, save the returned task_id to query progress and retrieve the final media URL.
  • Generated clips are delivered as standard video files ready for playback and editing software.

Kling 3.0 Turbo Standard Text to Video API frequently asked questions

What is the Kling 3.0 Turbo Standard Text to Video API?

Kling 3.0 Turbo Standard Text to Video is a Kuaishou Kling AI model for generating video from text prompts. It creates 720p clips with high-throughput generation, native lip-synced multilingual audio, and multi_prompt storyboarding for up to 6 shots. Built on the Kling 3.0 Turbo high-speed generation stack, it follows prompt narrative and camera direction while keeping character appearance, scene continuity, and audiovisual alignment across cuts. You can call it programmatically or try it from the playground above. Without multi_prompt, duration is 3–15 seconds. With multi_prompt, omitting top-level duration also allows a shot total of 1 or 2 seconds.

How does Kling 3.0 Turbo Standard Text to Video use multi_prompt?

multi_prompt accepts 1–6 shots. Each shot needs a nonblank prompt and an optional integer duration of 1–15 seconds, defaulting to 5. The sum must not exceed 15 seconds. Omit top-level duration to use that sum, including totals of 1 or 2 seconds. If supplied, top-level duration must be an integer from 3 to 15 and equal the sum. multi_prompt cannot be combined with a nonblank top-level prompt. Edit storyboards in JSON mode; switching to the form keeps JSON mode and all shot settings.

How long can Kling 3.0 Turbo Standard Text to Video generate?

Without multi_prompt, duration is an integer from 3 to 15 seconds and defaults to 5. With multi_prompt, each shot defaults to 5 seconds and the sum must not exceed 15. Omit top-level duration to use the sum, including totals of 1 or 2 seconds; an explicit top-level duration must be 3–15 and equal that sum.

Does Kling 3.0 Turbo Standard Text to Video generate native audio?

Yes. Native audio is generated with the picture, with lip sync aligned to multilingual speech plus matching ambience and action cues. Add dialogue language, tone, or environmental sound directions in the prompt to guide the soundtrack; there is no separate sound toggle field.

Which aspect ratios does Kling 3.0 Turbo Standard Text to Video support?

It supports 16:9, 9:16, and 1:1, with 16:9 preselected in the Playground. Use 16:9 for landscape storytelling, 9:16 for vertical shorts, and 1:1 for square layouts; full options are listed in the Parameters section.

How does Kling 3.0 Turbo Standard Text to Video differ from Pro?

Standard is fixed at 720p and billed at 17 credits per second—best for high-throughput drafts and batch iteration. The sibling Pro Text to Video tier is fixed at 1080p for sharper finals. Both support native audiovisual sync and multi_prompt storyboarding, with 3–15 second durations or shot totals of 1–2 seconds when top-level duration is omitted; choose by clarity needs versus cost.

How is Kling 3.0 Turbo Standard Text to Video billed per second?

Billing is by output seconds at 17 credits/sec for 720p ($0.085/sec), where 1 credit = $0.005. A default 5-second run costs 85 credits ($0.425); longer clips cost more—see the Pricing section for the exact rate.