Kling 1.6 Standard Text to Video API

kwaivgi/kling-video/v1.6/standard/text-to-video

Kling 1.6 Standard Text to Video transforms natural language prompts into cohesive 720p video clips, with flexible 5-second or 10-second durations, realistic physical dynamics, and cinematic camera control. It preserves prompt intent and spatial continuity across dynamic scene transitions while rendering authentic lighting and natural character movement.

Input
548/2500
78/2500
OutputReady
9 credits / second × 5 seconds = 45 credits · $0.225
Continue with

Examples

A continuous five-second vertical cinematic shot of one small unbranded red alpine cable-car cabin suspended from a taut overhead cable. The camera follows slowly alongside the cabin as it emerges from a thin bank of white cloud; distant rugged snow peaks gradually appear behind it. The cabin maintains its rigid shape, hangs upright and advances steadily along the cable, with subtle realistic sway. Cold blue shadows and warm early sun. Spacious mountain atmosphere, coherent depth and restrained camera movement, no cuts, no lettering or logos.

Slow-motion wildlife footage of a soaking-wet capybara vigorously shaking water off on a muddy riverbank. The action begins immediately: its head swings visibly from side to side, followed by a rippling shake through the shoulders and wet fur. A clearly visible spray of bright water droplets bursts outward and falls around the animal. Capture this single shake over the full five-second shot, medium close side view, warm backlighting, stationary camera. One capybara with coherent anatomy, realistic fur and water. No cuts, no other animals, no text or logos.

One continuous five-second hand-drawn 2D animation with watercolor textures on a quiet peach-colored beach. Exactly one small rust-red hermit crab, with two eye stalks and a large pale spiral shell firmly on its back, takes several short sideways steps to the right across damp sand. Its oversized shell gently rocks with each step while its eyes remain curious and alert. Simple consistent anatomy, weighty shell, tiny footprints, turquoise sea in the distance. Static full-body side view, warm illustrated natural-history charm. No transformation, no extra characters, no cuts, no text, no logos.

Kling 1.6 Standard Text to Video

Kling 1.6 Standard Text to Video is developed by Kuaishou (Kwaivgi) for cost-effective, high-quality text-to-video workflows. Developers and creators can submit descriptive prompts up to 2,500 Unicode characters to choreograph multi-stage subject action, ambient lighting, and expressive camera movement. The model renders continuous 5-second or 10-second videos in 720p resolution with physically believable motion, flexible aspect ratio selection, and transparent 9 credits per second pricing.

Why Choose This?

  • Direct Text-to-Video GenerationBring creative visions to life straight from descriptive prompts without requiring initial image assets or storyboards.

  • Realistic Physical DynamicsSimulates real-world gravity, fluid movements, and fabric inertia for authentic character movement and environmental interaction.

  • Multiple Aspect RatiosSupports 16:9 widescreen, 9:16 vertical, and 1:1 square compositions for seamless deployment across desktop, mobile, and social platforms.

  • Fine-Tuned Prompt GuidanceSupports up to 2,500 characters, optional negative prompts, and cfg_scale tuning from 0 to 1 to balance prompt adherence with creativity.

  • Predictable Per-Second PricingBilled at an affordable 9 credits per second ($0.045/s)—45 credits for 5s and 90 credits for 10s—with automatic credit refunds on failed tasks.

Parameters

ParameterRequirementDescription
promptRequired

Required nonblank string, at most 2,500 Unicode characters after trimming.

durationRequired

Required integer: 5 or 10 seconds. No strings, booleans or fractional durations. No API default; the playground starts at 5 seconds.

aspect_ratioOptional

Optional: 1:1, 16:9 or 9:16. No default.

negative_promptOptional

Optional string, at most 2,500 Unicode characters.

cfg_scaleOptional

Optional finite number from 0 to 1. No default. Not supported in Elements.

How to Use

  1. Describe Subject and SettingDefine character appearances, primary subjects, and lighting conditions in the opening sentences of your prompt.

  2. Specify Camera MovementIncorporate cinematic camera language such as pan, tilt, zoom, or tracking shots to establish depth and perspective.

  3. Configure Duration and RatioSelect 5 seconds or 10 seconds and choose an aspect ratio (16:9, 9:16, or 1:1) that fits your destination platform.

  4. Adjust Negative Prompt and GuidanceOptionally add negative_prompt to suppress unwanted artifacts and tune cfg_scale to control prompt adherence.

  5. Submit and Retrieve OutputPost your request to the asynchronous submit endpoint and poll with task_id to download the completed MP4 video.

Pricing

9 credits / second · $0.045 / second. 1 credit = $0.005. Fal comparison: $0.056 / second; save 20%.

UsageRateDetails
5 seconds45 credits · $0.2259 credits × 5 seconds
10 seconds90 credits · $0.4509 credits × 10 seconds

Best Use Cases

  • Social Media and Marketing ClipsQuickly produce engaging, high-impact video content for TikTok, Instagram Reels, and digital campaigns.

  • Creative Film PrevisualizationPrototype cinematic shot sequences, pacing, and camera angles during early pre-production to test script concepts.

  • Advertising Concept MockupsConvert storyboard drafts and copy ideas into dynamic video pitches with believable lighting and movement.

  • Game and Animation VisualsVisualize fantasy and sci-fi environments, atmospheric landscapes, and dynamic character introductions.

Pro Tips

  • Use Action-Oriented Verbs: Structure prompts with sequential action phrases like 'walks toward the camera and turns around' for smoother motion.
  • Utilize the 2,500-Character Capacity: Add rich details about lighting, textures, and ambient dynamics (e.g., 'falling leaves drifting in the wind') to enhance visual depth.
  • Calibrate cfg_scale for Your Goal: Set cfg_scale to 0.3–0.5 for stylistic creative freedom, or 0.7–1.0 for strict adherence to descriptive scripts.
  • Filter Artifacts with Negative Prompts: Add terms such as 'blurry, distorted anatomy, overexposed, low quality' to maintain clean output.
  • Match Duration to Scene Complexity: Choose 5 seconds (45 credits) for quick cutaways and 10 seconds (90 credits) for developing narrative arcs.

Notes

  • Text-Only Input Contract: This endpoint only accepts prompt strings and text control parameters; image input fields are not supported.
  • 720p Resolution Output: Standard tier renders in 720p resolution; choose the Pro endpoint if your workflow requires native 1080p full HD.
  • Asynchronous Task Execution: Tasks are tracked asynchronously via unique task_id; credits are deducted upon validation and refunded if generation fails.

Kling 1.6 Standard Text to Video API frequently asked questions

What is the Kling 1.6 Standard Text to Video API?

Kling 1.6 Standard Text to Video is a Kuaishou (Kwaivgi) model for generating video clips from text prompts. It produces 720p resolution videos from natural language descriptions with 5-second and 10-second duration options, realistic physical dynamics, and multiple cinematic aspect ratios. Built on advanced multimodal diffusion architecture, it preserves scene composition and lighting continuity while rendering fluid, natural motion. You can call it programmatically or try it from the playground above.

Does Kling 1.6 Standard Text to Video support 10-second video generation?

Yes. The model provides both 5-second and 10-second duration options. Generating a 10-second clip costs 90 credits and enables longer continuous action, multi-stage storytelling, and gradual camera transitions.

What aspect ratios are supported by Kling 1.6 Standard Text to Video?

The endpoint supports three standard aspect ratios: 16:9 widescreen, 9:16 vertical, and 1:1 square. You can configure aspect_ratio in your request to match desktop displays, mobile feeds, or square banners.

What does the cfg_scale parameter control in Kling 1.6 Standard Text to Video?

The cfg_scale parameter controls prompt adherence as a finite number between 0 and 1. Higher values force the generation to follow the text prompt strictly, while lower values give the model greater creative flexibility.

How is Kling 1.6 Standard Text to Video priced?

The endpoint is billed per second at 9 credits per second (1 credit = $0.005, or $0.045 per second). A 5-second clip costs 45 credits ($0.225) and a 10-second clip costs 90 credits ($0.450), with automatic refunds if generation fails.

What is the maximum prompt length for Kling 1.6 Standard Text to Video?

The prompt field supports up to 2,500 Unicode characters after trimming whitespace. This generous limit allows for detailed descriptions of characters, environments, actions, and camera movements.

When should I choose Kling 1.6 Standard instead of Pro?

Standard is ideal for cost-sensitive workflows, early concept exploration, and high-volume iterations at 720p resolution for 9 credits per second. Choose the Pro endpoint when your project requires native 1080p full HD resolution and enhanced fine detail.