Kling 2.6 Pro Text to Video API

kwaivgi/kling-v2.6-pro/text-to-video

Kling 2.6 Pro Text to Video transforms text prompts into 1080p cinematic videos, featuring native audiovisual synchronization, flexible 5s and 10s durations, and industry-standard aspect ratios. It adheres faithfully to complex scene directions and real-world physical dynamics while preserving spatiotemporal coherence and crisp detail stability.

Input

747/1,000

Output
Ready
5 s · With audio · 120 credits ($0.600) / video
Continue using

Examples

One continuous five-second square cinematic close-up inside an old cinema projection booth. A working metal 35mm film projector fills the frame: visible upper reel turning steadily and a taut strip of film advancing through its gate. A brilliant warm projection beam cuts diagonally through dusty darkness toward the unseen screen. Slow subtle push-in, tangible worn metal, warm amber light, drifting dust, coherent mechanical motion. Native synchronized sound: steady rhythmic projector clatter and a gentle motor whirr, with the enclosed booth's slight resonance. No dialogue, no music, no people, no labels, readable writing, logos, commercial product staging, advertising or cuts.

One continuous five-second vertical sports documentary shot inside a bright empty squash court. A single adult player in plain unbranded white sportswear is seen from a rear three-quarter angle, full body and front wall visible. The player makes one controlled forehand strike with a squash racket; one small black ball flies to the front wall, visibly bounces once and returns low toward the floor. The player settles their stance. Fixed camera, anatomically natural movement, believable racket contact and one coherent ball trajectory. Native synchronized audio: a sharp racket tap, a hollow wall impact and a brief shoe squeak with clear indoor court echo. No crowd, commentary, music, slow motion, extra players, writing, branding or cuts.

Kling 2.6 Pro Text to Video Overview

Kling 2.6 Pro Text to Video is an advanced text-to-video generation model developed by Kuaishou Technology. It renders detailed text prompts directly into 1080p full high-definition video assets, features single-pass native synchronized audio synthesis, and accommodates 5-second or 10-second durations across 16:9, 9:16, and 1:1 aspect ratios.

Why Choose Kling 2.6 Pro Text to Video

  • 1080p Cinematic ResolutionProduces pristine 1080p full HD visuals with rich micro-expressions, refined lighting shifts, and tactile environmental textures.

  • Native Audio SynchronizationGenerates synchronized ambient soundscapes and action-aligned Foley sound effects in a single inference step without external dubbing.

  • Accurate Physical DynamicsFaithfully simulates real-world kinetics, cloth draping, and fluid movements to ensure sweeping motions remain organic and believable.

  • Multi-Ratio & Flexible LengthsDelivers cohesive 5-second or 10-second clips across widescreen 16:9, vertical 9:16, and square 1:1 display formats.

  • Transparent Pricing & Automatic RefundsFeatures predictable credit pricing across silent and audio tiers, with automatic refunds if a generation task encounters an error.

Parameters

ParameterRequirementDescription
promptYes

Required nonblank string, 1 to 1,000 characters after trimming leading and trailing whitespace.

Default-
durationYes

Required integer: 5 or 10 seconds. Strings and fractional durations are rejected.

Default-
aspect_ratioYes

Required: 16:9, 9:16 or 1:1. Controls the output video aspect ratio.

Default-
soundYes

Required boolean: true for audio, false for no audio.

Default-

How to Use

  1. Define scene and cinematographyDetail subject appearance, sequential actions, ambient lighting, and camera trajectories within your text prompt (up to 1,000 characters).

  2. Select duration and aspect ratioChoose between 5s or 10s output lengths, and set the aspect ratio (16:9, 9:16, or 1:1) matching your target distribution medium.

  3. Configure synchronized audioEnable sound to synthesize matching environmental noise and movement audio, or leave sound disabled for purely visual output.

  4. Dispatch generation jobSubmit your API request to receive a unique task ID, initiating background execution in the processing pipeline.

  5. Retrieve final videoPoll the task status endpoint until finished, then access the verified 1080p MP4 download URL.

Pricing

Billed per generated video. 1 credit = $0.005.

UsageRateDetails
5 s · No audio65 credits/video$0.325/video
10 s · No audio130 credits/video$0.650/video
5 s · With audio120 credits/video$0.600/video
10 s · With audio240 credits/video$1.20/video

Best Use Cases

  • Cinematic Storyboard PrevisualizationConvert screenplay descriptions into dynamic 1080p visual animatics to evaluate pacing and lighting setups before production.

  • High-Impact Social Video AdsCraft mobile-native 9:16 short clips with synchronized sound effects designed to captivate audiences across social feeds.

  • E-Commerce Dynamic Concept VisualsDescribe product demonstrations and lifestyle environments to produce high-end commercial hero videos.

  • Digital Creative Concept ExplorationTranslate surreal imaginative concepts into moving visual sequences guided by natural physical laws.

Pro Tips

  • Structure prompts hierarchically: core subject appearance first, followed by chronological actions, environmental context, camera movements, and lighting mood.
  • Utilize standard cinematography terms such as slow push-in, low-angle pan right, or steady tracking shot for smooth camera motion.
  • When enabling sound, incorporate descriptive audio cues (such as footsteps splashing on pavement or birds chirping in the morning mist) to guide acoustic synthesis.
  • Ensure complex physical actions follow realistic momentum progression rather than specifying instantaneous opposing movements.

Notes

  • Prompt is required, accepting between 1 and 1,000 characters after trimming whitespace.
  • Duration accepts integer values of 5 or 10 seconds; other values will be rejected during validation.
  • Generation operates asynchronously: requests return a task_id immediately, with status retrievable via polling or callback_url webhooks.

Kling 2.6 Pro Text to Video API Frequently Asked Questions

What is the Kling 2.6 Pro Text to Video API?

Kling 2.6 Pro Text to Video is a Kuaishou Technology model for generating cinematic video from text prompts. It creates 1080p full high-definition videos in 5-second or 10-second durations, featuring native audiovisual synchronization and flexible framing across 16:9, 9:16, and 1:1 aspect ratios. Built on advanced spatiotemporal diffusion architectures, it preserves narrative continuity and realistic physical dynamics while rendering rich visual detail. You can call it programmatically or try it from the playground above.

What types of audio does Kling 2.6 Pro Text to Video generate?

When sound is enabled, the model synthesizes context-aware audio tracks directly during inference, including ambient soundscapes, physical action sound effects, and situational Foley noise, removing the necessity of manual post-production audio editing.

Does Kling 2.6 Pro Text to Video support 10-second generations?

Yes. The endpoint supports both 5-second and 10-second output options. Selecting 10 seconds provides extended narrative continuity, elaborate character motion sequences, and sustained atmospheric flow.

Which aspect ratios are supported by Kling 2.6 Pro Text to Video?

It supports 16:9 widescreen, 9:16 vertical, and 1:1 square aspect ratios. The model renders frames natively at the selected ratio, eliminating the composition loss associated with post-generation cropping.

How is Kling 2.6 Pro Text to Video billed?

Pricing is determined by duration and audio selection: 65 credits ($0.325) for 5s without sound, 130 credits ($0.650) for 10s without sound, 120 credits ($0.600) for 5s with sound, and 240 credits ($1.200) for 10s with sound. If a task terminates due to a system error, reserved credits are automatically refunded in full.

How does Kling 2.6 Pro Text to Video maintain motion continuity?

Through deep spatiotemporal attention modeling, the architecture maintains cohesive motion trajectories across frames, preserving consistent human anatomy and environmental spatial geometry during dynamic action sequences.

What are the best practices for writing Kling 2.6 Pro Text to Video prompts?

Organize prompts into distinct layers covering character identity, specific physical actions, surrounding environment, camera trajectory, and lighting atmosphere to provide precise composition anchors for the diffusion process.