Wan 2.7 Image Text to Image API

alibaba/wan-2.7/text-to-image
6 preset sizes / custom size · 1–4 images

Wan 2.7 Image Text to Image transforms natural language prompts into high-fidelity 1K and 2K images, with Thinking Mode spatial reasoning, bilingual typography rendering, and palette-guided color control. It strictly adheres to prompt physical laws and spatial layouts while rendering authentic lighting textures and rich details.

Get API Key
Input
716/5000
OutputReady
1 × 4.2 = 4.2 credits · $0.021

Continue with

Examples

standard-text-01-output

Photorealistic quiet old railway waiting room, warm late-afternoon sunlight, worn wooden benches and cream plaster. A large dark green timetable board is the main focal point, nearly front-facing and occupies half the frame. Render ONLY these five lines of crisp, large, correctly spelled lettering on the board: "列车时刻 / TIMETABLE", "松林 / PINES 08:20", "湖畔 / LAKESIDE 10:45", "山谷 / VALLEY 16:30", "一路平安 / SAFE JOURNEY". Carefully aligned rows with consistent spacing, all lettering readable. Authentic painted enamel, subtle scratches, dust in the light, cinematic but restrained composition. No other visible text, people, promotional graphics or extra boards. No advertising, brands, logos, watermark, or foxes.

standard-text-02-output

Natural wildlife photograph at eye level of one roseate spoonbill foraging in a salt marsh. Its long flattened spoon-shaped bill touches the shallow water; anatomically correct slender legs and layered pale pink feathers, subtle crimson shoulder patch. One bird only, no duplicates. Pale salt grasses, small ripples and a coherent reflection, layered muted blue-green estuary in the distance. Soft overcast morning light and delicate feather detail, unposed natural behavior. No writing. No advertising, brands, logos, watermark, or foxes.

standard-text-03-output

Candid documentary portrait of a silver-haired woman in her seventies considering a chess move at an outdoor stone chess table in a leafy neighborhood square. Distinctive deep smile lines, short wavy white hair, ochre cardigan, natural skin texture. Her right hand gently holds one wooden knight between thumb and index finger above the board; her left rests on the edge of the table. Believable fingers and wrists, restrained thoughtful expression. Green plane-tree shade with soft sunlight, intimate waist-up composition, no other people or text. No advertising, brands, logos, watermark, or foxes.

Wan 2.7 Image Text to Image Overview

Wan 2.7 Image Text to Image is a high-performance text-to-image model developed by Alibaba Tongyi Lab. Built on a unified multimodal planner and diffusion transformer (DiT) architecture, it analyzes multi-object layout and prompt nuances through an integrated chain-of-thought phase prior to synthesis. The model supports 6 preset aspect ratios as well as custom pixel dimensions, excelling in multi-line typography, brand color accuracy, and realistic portrait skin rendering for high-volume concept art, advertising, and marketing assets.

Why Choose

  • Thinking Mode Spatial ReasoningAnalyzes multi-object relationships, depth layers, and occlusion constraints during a built-in reasoning stage before pixel synthesis, ensuring coherent compositions for complex prompts.

  • Bilingual In-Image TypographyEliminates distorted AI lettering by directly rendering sharp, legible Chinese and English text, paragraphs, and slogans with crisp typographic outlines.

  • Palette-Guided Color & Extreme Aspect RatiosAccepts hex color arrays for precise brand palette adherence and natively accommodates extreme 1:8 to 8:1 aspect ratios for panoramic banners and vertical scrolls.

Parameters

ParameterRequirementDescription
promptRequired

String, 1–5,000 characters after trimming.

sizeOptional

Six preset sizes, or an object such as {"width":1280,"height":720}. Width and height must be positive integers.

Default1024x1024512x512768x10241024x768576x10241024x576
nOptional

Output image count, integer 1–4. Credits = per-image rate × n.

Default1
seedOptional

Optional integer, omitted when unset. The upstream documentation specifies no numeric range.

How to Use

  1. Compose Prompt and Typography TextDescribe visual subjects and lighting in detail, enclosing any desired banner text in quotation marks and specifying color preferences.

  2. Select Dimensions and CountChoose a preset aspect ratio or enter custom dimensions, then specify output count from 1 to 4.

  3. Submit Task and Fetch OutputsSubmit request to the unified API endpoint, track task_id status asynchronously, and retrieve final image URLs.

Pricing

4.2 credits per output image ($0.021). Total credits = 4.2 × n. Size and reference images add no charges. 1 credit = $0.005.

UsageRateDetails
Standard4.2 credits/image · $0.021Multiply by output count n; credits are refunded for failed tasks.

Use Cases

  • Commercial Ads & Social Media BannersRender crisp promotional text and brand palettes directly inside high-resolution marketing visuals.

  • E-Commerce Concept Sets & StagingPlan spatial placement and realistic light reflections for believable merchandise settings and editorial lookbooks.

  • Panoramic Layouts & Digital ArtLeverage extreme 1:8 to 8:1 ratios to design website hero headers, long infographic scrolls, and visual storytelling boards.

Tips

  • Optimize Typography Rendering: Place text phrases inside quotation marks (e.g., a billboard reading "SUMMER SALE") and state font style and placement clearly.
  • Activate Spatial Reasoning: Specify spatial positions like "in the foreground left" or "behind the counter" to help the reasoning planner allocate scene composition.
  • Maintain Visual Series Consistency: Fix the seed number across prompts with matching lighting terms to generate cohesive visual asset packs.

Notes

  • Text-Only Input: This endpoint accepts text prompts only; use Wan 2.7 Image Edit when reference images are needed.
  • Character Count Boundary: Prompts must contain between 1 and 5,000 characters after whitespace trimming.
  • Output Quantity Range: Parameter n accepts integers 1 to 4, with billing scaling linearly per generated image.

Related Models

Wan 2.7 Image Text to Image API FAQ

What is the Wan 2.7 Image Text to Image API?

Wan 2.7 Image Text to Image is an Alibaba Tongyi Lab model for high-fidelity image generation from text prompts. It produces sharp 1K and 2K visuals with Thinking Mode spatial reasoning, bilingual typography rendering, and palette-guided color control. Built on a unified multimodal planner and DiT architecture, it preserves coherent physical laws and spatial arrangements while delivering authentic lighting textures. You can call it programmatically or try it from the playground above.

Can Wan 2.7 Image Text to Image generate multiple images in one request?

Yes. Setting parameter n between 1 and 4 outputs up to four distinct visual variants simultaneously. Each image maintains full quality and resolution, billed at 4.2 credits per output image.

Does Wan 2.7 Image Text to Image render legible text in images?

Yes. The model incorporates specialized multilingual glyph comprehension, rendering sharp Chinese and English words, titles, and signage directly inside the scene when enclosed in quotation marks.

How do I customize aspect ratio and resolution in Wan 2.7 Image Text to Image?

You can select from 6 preset options (including 1:1, 16:9, 9:16, 4:3, and 3:4) or pass a custom {width, height} object with positive integer pixel values, spanning ratios from 1:8 to 8:1.

How does Thinking Mode work in Wan 2.7 Image Text to Image?

When processing dense or multi-subject prompts, an internal reasoning stage analyzes spatial arrangements, perspective depth, and object interactions before diffusion begins to guarantee composition balance.

Can I use a fixed seed to reproduce images with Wan 2.7 Image Text to Image?

Yes. Providing an integer in the optional seed parameter preserves the initial latent noise layout, allowing you to fine-tune prompts while maintaining consistent composition structures.

How does billing work if a generation task fails?

Vidgo provides an automatic credit refund guarantee. If a task fails due to validation errors, network interruption, or processing timeout, all deducted credits are immediately returned to your account balance.