Kling O3 Image Text to Image API

kwaivgi/kling-image-o3/text-to-image
1K / 2K / 4K · 1–9 images

Kling O3 Image Text to Image transforms text prompts into native high-resolution images up to 4K, featuring micro-level physical realism, reusable element identity locking, and 8 standard aspect ratios. It adheres faithfully to complex spatial narratives and illumination logic while maintaining rigorous aesthetic continuity across multi-image generation batches.

Get API Key
Input
1182/2000
Elements

Reference subjects in order with @Element1, @Element2. Provide a main view and additional views. Sign in to upload files.

OutputReady
1 × 3.5 = 3.5 credits · $0.018

Continue with

Examples

text-01-output.png

Create a warm storybook cinematic scene featuring one original little library robot: a squat ivory enamel spherical body, a teal horizontal glass faceplate with two amber circular eyes, an orange rectangular chest compartment, short brass articulated arms ending in simple grippers, one short antenna topped with an orange ball, and two stubby teal feet. Give its tactile painted-metal surfaces subtle scuffs and believable mechanical joints. In a circular old library with curved wooden shelves, the tiny robot is standing on the tips of its feet on a low wooden step, using both grippers to slide one oversized deep-blue clothbound book into a low shelf just above its head. The book is broad relative to the robot but physically manageable. A few books already fill that shelf. Full-body three-quarter view, robot in the left foreground, a curved balcony and spiral stair receding to the right, golden afternoon window light and floating dust, tactile miniature materials, charming believable mechanical pose, coherent perspective. One robot, two arms, two legs, one blue book being handled. No people, animals, visible writing, labels, logos, commercial messaging or watermarks.

text-02-output.png

纪实摄影,一位白发苍苍的中国老皮影艺人,穿着褪色的靛蓝棉布外套,在乡村小戏台后台的旧木桌前,用细针修补一件传统中国皮影戏偶的肩部连接处。皮影必须是薄而平的半透明染色牛皮剪影,侧面脸谱,红、赭黄和绿色的精细镂空花纹,关节用细线连接,配细竹签操纵杆;绝不是立体木偶或布偶。戏偶平放在桌面上,左手轻轻固定,右手持针。背景挂着三件同类平面彩色镂空皮影。侧窗自然光穿透皮革,呈现微微透光的暖色,真实的皱纹、双手、旧木纹和衣料质感。中景三分之四视角,老人面孔和完整双手与桌面都清晰可见,安静专注的开演前修补瞬间。只有一位老人,无其他人物,无广告、品牌、水印或文字。

text-03-output.png

Natural underwater wildlife photography of exactly one green sea turtle swimming calmly through shafts of sunlight in clear shallow tropical seawater. Three-quarter side view with the head facing right, the entire animal and all anatomically visible flippers within frame, realistic shell scutes and mottled skin, no human-like expression. The turtle occupies the middle left and glides toward open blue water on the right. Pale rippled sand and sparse seagrass well below it, a small coral outcrop confined to the lower left, shimmering surface near the top edge. Physically convincing light scattering and caustic patterns across the shell and sand, natural turquoise-to-blue depth gradient, crisp turtle detail with gently receding background. A tranquil observation of wildlife, no fantasy glow. No other animals, people, divers, buildings, plastic, typography, logos, frame or watermark.

Kling O3 Image Text to Image Overview

Kling O3 Image Text to Image is Kuaishou's (Kling AI) next-generation flagship multimodal text-to-image synthesis model. Built on advanced visual reasoning architectures and an audiovisual-inspired narrative aesthetic engine, it produces direct 1K, 2K, and 4K outputs without external upscaling. It natively supports 8 aspect ratios, parallel batch generation up to 9 images, and persistent element identity locking for commercial creative workflows.

Why Choose Kling O3 Image Text to Image?

  • Direct Native 4K ResolutionBypasses conventional post-generation upscaling by generating genuine 4K resolution natively, capturing skin pores, fabric textures, and subtle light reflections directly.

  • Element Consistency LockingAllows developers to register character, product, or branded assets via subject reference slots and bind them in prompts using @Element1 for consistent multi-generation assets.

  • Cinematic Narrative Aesthetic EngineDeconstructs director-level cinematography, rendering images with spatial depth, volumetric atmospheric lighting, and evocative visual mood.

  • Concurrent Batch Generation up to 9 ImagesConfigures single requests to produce between 1 and 9 distinct compositional variations, drastically reducing review turnaround times.

  • Predictable Tier Billing & Auto RefundsCharges transparently per output image according to clear resolution tiers with zero hidden fees, backed by automated credit refunds on task errors.

Parameters

ParameterRequirementDescription
promptRequired

String, 1–2,000 characters after trimming.

resolutionOptional

Output resolution determines the per-image rate.

Default1K2K4K
sizeOptional

Aspect ratio; auto is supported only by Edit. 16:9, 9:16, 1:1, 4:3, 3:4, 3:2, 2:3, 21:9

Default16:99:161:14:33:43:22:321:9
output_formatOptional

Output image format.

Defaultpngjpegwebp
nOptional

Integer output count, 1–9. Total credits = the resolution rate × n.

Default1
elementsOptional

Array of subject-reference objects. Reference them as @Element1, @Element2, etc. No published count limit; extension fields are preserved.

elements[].frontal_image_urlOptional

Public HTTP(S) URL for the main subject view.

elements[].reference_image_urlsOptional

Array of public HTTP(S) image URLs for additional views of the same subject.

How to Use

  1. Formulate a structured promptDescribe core subjects, spatial staging, lighting atmosphere, and visual style (up to 2,000 characters). Reference registered subjects using @Element1 where needed.

  2. Define subject elements (optional)Supply front-facing or multi-angle image URLs in the elements array to anchor recurring visual identities.

  3. Select resolution and aspect ratioChoose 1K, 2K, or 4K resolution, and pick an aspect ratio suited to your publication medium (such as widescreen 16:9 or vertical 9:16).

  4. Configure format and countSpecify target file format (JPEG, PNG, or WebP) and set the batch size n between 1 and 9.

  5. Dispatch request and retrieve assetsSubmit the API job to receive a unique task ID, poll until completion, and download verified high-resolution images.

Pricing

1K/2K: 3.5 credits/image (approximately $0.018); 4K: 7 credits/image ($0.035). Total credits = resolution rate × n, without input-image or subject-reference surcharges. At $0.005 per credit, 3.5 credits equals $0.0175. USD is rounded to three decimals after calculating the total.

UsageRateDetails
1K / 2K3.5 credits/image · ≈ $0.018Multiply by the output count n.
4K7 credits/image · $0.035Multiply by the output count n.

Best Use Cases

  • Commercial advertising campaigns and key visualsLeverage native 4K output to produce billboard-ready marketing graphics, digital campaigns, and print-grade posters.

  • Cinematic concept art and production storyboardingTranslate detailed narrative scripts into atmosphere-rich concept paintings with convincing physical illumination.

  • Digital character and mascot asset developmentUtilize element identity locking to generate recurring characters across diverse postures, wardrobe variations, and scene environments.

  • E-commerce merchandise and lifestyle stagingDescribe product materials and studio setups to generate photo-studio-quality packshots and ambient retail visuals.

Pro Tips

  • Structure prompts hierarchically: core subject appearance, pose, spatial perspective, lens specifications, lighting conditions, and material textures.
  • To achieve authentic photographic fidelity, incorporate specific optical parameters like '85mm portrait lens', 'shallow depth of field', 'subtle rim lighting', or 'soft diffused studio illumination'.
  • When maintaining a persistent character identity, supply a high-clarity frontal view via elements and explicitly tag the subject with @Element1 in the prompt.
  • Rapidly iterate concept compositions at 1K resolution first, then switch to native 4K for final commercial deliverables.

Notes

  • The prompt parameter is required and must contain 1 to 2,000 characters after trimming whitespace.
  • Output count n must be an integer between 1 and 9, inclusive.
  • This text-to-image endpoint rejects image_urls; use Kling O3 Image Edit when reference images are required.
  • Generation is asynchronous; track progress via task status polling or configure a callback_url for automatic completion notifications.

Related Models

Kling O3 Image Text to Image API frequently asked questions

What is the Kling O3 Image Text to Image API?

Kling O3 Image Text to Image is a state-of-the-art multimodal text-to-image synthesis model developed by Kuaishou (Kling AI). It transforms descriptive text prompts into native high-resolution images up to 4K, featuring micro-level tactile realism, element identity locking, and support for 8 standard aspect ratios. Engineered with unified visual intelligence and narrative aesthetic mechanisms, it preserves complex spatial logic and illumination physics while rendering rich, cinematic visual quality. You can call it programmatically or try it from the playground above.

Do 4K images generated by Kling O3 Image Text to Image require upscaling?

No. Kling O3 Image renders 4K resolution natively, synthesizing authentic skin pores, fine fabric weaves, and delicate light gradients directly rather than interpolating or stretching lower-resolution previews.

How does Kling O3 Image Text to Image maintain subject consistency?

By supplying frontal and reference image URLs within the elements array and citing them as @Element1 or @Element2 in your prompt, the model locks facial geometry, physique, and styling attributes across brand-new poses and backgrounds.

How many images can Kling O3 Image Text to Image generate in a single call?

A single API request supports generating between 1 and 9 independent image variants via the n parameter, enabling swift exploration and comparison under identical prompt settings.

What aspect ratios and output formats does Kling O3 Image Text to Image support?

It supports 8 framing formats: 16:9, 9:16, 1:1, 4:3, 3:4, 3:2, 2:3, and 21:9, with file format choices including JPEG, PNG, and WebP (PNG by default). Compositions and lighting are computed natively for each ratio.

How is Kling O3 Image Text to Image billed?

Pricing is calculated per output image based on resolution: 3.5 credits (approx. $0.018) for 1K or 2K, and 7 credits ($0.035) for 4K, multiplied by the image count n. There are no surcharges for input text or element objects, and failed jobs receive an immediate full credit refund.

What is the best way to structure prompts for Kling O3 Image Text to Image?

Organize prompts from primary subject traits to spatial depth, lighting mood, and physical textures. Specifying explicit spatial relationships and camera staging enables the engine to balance visual focal depth realistically.