Wan 2.7 Image Pro Text to Image API

alibaba/wan-2.7/text-to-image-pro
6 preset sizes / custom size · 1–4 images

Wan 2.7 Image Pro Text to Image transforms natural language prompts into native 4K commercial-grade images, delivering micro-skin textures, flagship Thinking Mode spatial reasoning, and expansive typography rendering. It strictly follows complex compositional prompts while rendering print-ready sharpness and rich dynamic range.

Get API Key
Input
771/5000
OutputReady
1 × 10.5 = 10.5 credits · $0.0525

Continue with

Examples

pro-text-01-output

Intimate documentary photograph of an adult Andean woman in her fifties working at a traditional wooden floor loom in a quiet adobe weaving room. Strong individual facial character, braided dark hair streaked with silver, subtle wrinkles, a simple indigo blouse. Both anatomically natural hands gently separate a small group of taut ivory warp threads, with clear separation between fingers and threads. Hundreds of fine parallel threads recede toward the loom; a partially woven terracotta and deep-blue geometric textile lies in the foreground. One broad skylight gives soft directional illumination across skin, worn wood and fine fibers. Medium-wide composition, tactile detail, realistic restrained color, no text. No advertising, brands, logos, watermark, or foxes.

pro-text-02-output

Still-life photograph on a matte pale stone floor in warm angled window light. A small open undyed linen drawstring pouch lies at the upper left. Exactly FOUR separate glass marbles have rolled out, arranged in one loose diagonal row: one red, one blue, one green, one amber. Exactly ONE old wooden spinning top lies on its side to their right, with its pointed tip visible. The four marbles do not overlap and are all entirely visible; no other balls, beads or round objects. Correct translucent colored glass, delicate caustics and consistent cast shadows, tactile linen and worn wood, generous negative space. No text. No advertising, brands, logos, watermark, or foxes.

pro-text-03-output

Wide natural landscape photograph of a geothermal pool on a remote volcanic highland immediately after dawn. Thin translucent steam rises over turquoise water; delicate white hoarfrost coats jagged dark basalt stones along the near shore. Farther away muted rust-red mineral terraces lead to snow-dusted volcanic slopes beneath a pale lavender sky. Carefully separated foreground crystal detail, midground steam and distant mountain layers, believable warm water beside frozen ground, soft amber side light, subtle reflected sky. No people, buildings, animals or writing. No advertising, brands, logos, watermark, or foxes.

Wan 2.7 Image Pro Text to Image Overview

Wan 2.7 Image Pro Text to Image is a flagship text-to-image model developed by Alibaba Tongyi Lab for commercial printing, brand design hubs, and professional digital artistry. Unlike standard tiers, the Pro edition synthesizes native 4K outputs without secondary upscaling artifacts. Featuring an advanced Thinking Mode reasoning engine, it orchestrates complex multi-subject perspective and layout logic while providing comprehensive support for bilingual typography and hex-color palette guidance, delivering industrial-grade visual assets for concept art, billboards, and luxury marketing.

Why Choose

  • Native 4K Commercial Super-ResolutionOutputs uncompressed 4K native resolutions directly, exposing pores, hair sheen, and fabric weaves to satisfy high-end editorial and print criteria.

  • Flagship Thinking Mode Spatial PlanningEnhanced chain-of-thought analysis resolves intricate multi-subject perspectives and layered lighting, minimizing distortions across dense descriptive prompts.

  • Professional Typography & Color PalettesEffortlessly renders multi-line paragraphs in Chinese and English while strictly aligning color proportions with specified hex palettes for brand consistency.

Parameters

ParameterRequirementDescription
promptRequired

String, 1–5,000 characters after trimming.

sizeOptional

Six preset sizes, or an object such as {"width":1280,"height":720}. Width and height must be positive integers.

Default1024x1024512x512768x10241024x768576x10241024x576
nOptional

Output image count, integer 1–4. Credits = per-image rate × n.

Default1
seedOptional

Optional integer, omitted when unset. The upstream documentation specifies no numeric range.

How to Use

  1. Write Detailed High-Fidelity PromptsDefine subjects, tactile surfaces (like silk or skin pores), in-image typography text, and hex color values in detail.

  2. Configure 4K Dimensions and QuantitySelect high-resolution presets or enter custom 4K pixel dimensions (e.g. 3840x2160), then set output count from 1 to 4.

  3. Submit Pro Request and Retrieve AssetsDispatch the task to the asynchronous endpoint, monitor progress until complete, and download uncompressed super-resolution images.

Pricing

10.5 credits per output image ($0.0525). Total credits = 10.5 × n. Size and reference images add no charges. 1 credit = $0.005.

UsageRateDetails
Pro10.5 credits/image · $0.0525Multiply by output count n; credits are refunded for failed tasks.

Use Cases

  • Commercial Billboards & Exhibition DisplaysOutput native 4K visuals suitable for large-format print media, product catalogs, and high-DPI digital signage.

  • Luxury Fashion & Cosmetic PhotographyPreserve translucent skin rendering, individual eyelashes, and micro-reflections for discerning beauty and luxury brand campaigns.

  • Cinematic Worldbuilding & Key ArtUtilize deep spatial reasoning to stage epic architectural vistas, complex light spill, and multi-character storytelling illustrations.

Tips

  • Describe Tactile Micro-Textures: Words like "individual hair strands", "translucent skin pores", and "diffuse ambient bounce" unlock the full rendering depth of the Pro model.
  • Leverage Multi-Line Typography: Enclose promotional headlines and copy inside quotation marks, clarifying alignment and hierarchy for crisp poster layouts.
  • Define Palette Hex Codes: Include exact hex values (e.g. "#1E3A8A occupying 60% of the canvas") to ensure outputs strictly match corporate identity guidelines.

Notes

  • Text-To-Image Exclusive: This endpoint generates images strictly from text; use Wan 2.7 Image Pro Edit when source images are required.
  • 4K Processing Time: Native 4K synthesis processes massive tensor volumes; generation times may slightly exceed standard 1K/2K jobs.
  • Billing Calculation: Each output image costs 10.5 credits, with total fees calculated as 10.5 × n and full refunds for failed tasks.

Related Models

Wan 2.7 Image Pro Text to Image API FAQ

What is the Wan 2.7 Image Pro Text to Image API?

Wan 2.7 Image Pro Text to Image is a flagship ultra-high-definition text-to-image model developed by Alibaba Tongyi Lab. It synthesizes native 4K commercial-grade visuals directly from text prompts, delivering micro-skin textures, advanced Thinking Mode spatial reasoning, and expansive typography rendering. Built on a unified multimodal planner and enhanced DiT architecture, it renders authentic physical lighting and print-ready fidelity across complex prompts. You can call it programmatically or try it from the playground above.

Is the 4K resolution in Wan 2.7 Image Pro Text to Image natively generated?

Yes. The Pro tier generates true native 4K pixel arrays during latent diffusion rather than utilizing post-process upscaling algorithms, ensuring authentic optical textures across hair, fabric, and focal backgrounds.

What are the primary differences between Wan 2.7 Image Pro and the Standard edition?

The Pro version unlocks native 4K resolution output and features a scaled-up Thinking Mode planner that excels in handling complex multi-subject interactions, intricate skin textures, and dense bilingual typography.

Does generating native 4K images cost extra credits in Wan 2.7 Image Pro Text to Image?

No. Wan 2.7 Image Pro Text to Image utilizes a flat rate of 10.5 credits ($0.0525) per output image regardless of whether you output 1K, 2K, or native 4K dimensions.

Does Wan 2.7 Image Pro Text to Image support multilingual typography?

Yes. The model demonstrates exceptional typographic comprehension across Chinese, English, and other languages, cleanly rendering headlines, subheadings, and body paragraphs with smooth vector-like font edges.

How can I generate extreme panoramic compositions with Wan 2.7 Image Pro Text to Image?

Specify custom positive pixel values in the size object parameter. The architecture accommodates aspect ratios up to 8:1 for wide panoramic illustrations and digital ribbon displays.

How does billing handle failed generations with Wan 2.7 Image Pro Text to Image?

Vidgo enforces an automatic refund policy. If an image generation job fails due to validation errors or system timeouts, the 10.5 credits deducted are restored to your balance immediately.