Z-Image Turbo API

alibaba/z-image/turbo

Z-Image Turbo transforms natural language prompts or reference images into high-quality visual assets with 8-step ultra-fast inference, native bilingual typography, and photorealistic rendering. It faithfully adheres to intricate prompt details and source composition while providing sub-second generation speed for high-volume creative iteration.

Input

0 / 1000

true
OutputWaiting

Your images will appear here

Submission charge: 2 credits · $0.010

Text-to-image and image editing: 2 credits / $0.010 per generation.

Examples

The Hidden Geometry of Red Cabbage

An intimate square macro photograph of a freshly cut red cabbage, its cross-section filling almost the entire frame. Intricate tightly folded violet and magenta leaves form an organic maze, separated by delicate ivory veins. A few tiny clear water droplets cling naturally to the cut surface. Soft daylight from the side reveals translucent leaf edges, crisp moist textures and gentle depth between the folds. Slightly oblique camera angle, a coherent botanical structure, sharply resolved central detail and subtle focus falloff at the outer edge. Quiet observational food photography, natural color, no arrangement of products, no packaging, no hands, no labels, no text, no logos and no watermark.

Warm Lights Along the Pool House

Add a few small wall-mounted lamps with a clearly visible warm amber glow beneath the eaves on the front facade of the existing low pool changing-room building. This is a subtle practical lighting addition to the existing neighborhood pool, not a redesign. Keep the daytime blue sky and natural colors. Preserve the principal camera viewpoint and layout: the rectangular swimming pool, tiled edge, dark lane lines, near-right steel ladder, low building on the left and background trees. The newly illuminated warm lights on the facade must be easy to see. Realistic architectural photograph, empty pool area, no people, no advertising, no text or watermark.

Z-Image Turbo

Z-Image Turbo is an efficient 6-billion-parameter (6B) lightweight diffusion transformer developed by Tongyi-MAI, engineered specifically for high-throughput, sub-second image generation and agile visual editing. Built on an innovative Single-Stream Diffusion Transformer (S3-DiT) architecture with latent flow-matching distillation, it completes high-fidelity synthesis in just 8 function evaluations (8 NFEs). The model excels in English and Chinese bilingual text rendering, natural photorealistic textures, and flexible aspect ratio adaptation. It also provides single-image editing via one reference image URL, backed by a commercial-friendly Apache 2.0 license and transparent flat pricing of 2 credits ($0.010) per run across both modes.

Why Choose Z-Image Turbo

  • 8-Step Ultra-Fast Inference & Sub-Second LatencyPowered by a Single-Stream Diffusion Transformer (S3-DiT) with latent distillation, it completes high-fidelity image rendering in just 8 denoising steps, achieving sub-second response times on enterprise GPUs.

  • Accurate Bilingual Typography in English & ChineseDeeply optimized for in-image typography, rendering sharp, legible characters on signage, packaging, and banners when target words are specified in prompts.

  • Photorealistic Details & Natural Studio LightingExcels at delicate skin textures, micro-expressions, macro product reflections, and authentic depth of field comparable to large-scale commercial generation models.

  • Dual-Mode Text-to-Image & Single-Image EditingSupports both direct prompt-to-image creation and targeted editing from a single reference image URL, accommodating fresh ideation and asset modification alike.

  • Apache 2.0 Commercial License & Predictable PricingReleased under the open Apache 2.0 license for unrestricted commercial deployments, paired with flat transparent pricing of 2 credits ($0.010) per generation.

Parameters

ParameterRequirementDescription
promptRequired

Describe the image to generate or the edits to apply. Use 1–1000 characters, including at least one non-whitespace character.

image_urlsOptional

Providing this field selects image editing. The array contains one HTTP(S) image URL.

sizeRequired for text-to-image; optional for editing

Output aspect ratio. Provide for text-to-image; optional for image editing.

1:14:33:416:99:16
enable_safety_checkerOptional

Boolean switch that controls the safety checker.

Defaulttruefalse

How to Use

  1. Select Generation Mode & InputsInput your text prompt and aspect ratio for standard text-to-image creation, or provide a single public HTTP(S) image URL in image_urls for single-image editing.

  2. Craft Descriptive Prompts & TextSpecify subjects, lighting directions, spatial framing, and materials in detail. Wrap desired in-image English or Chinese text in quotation marks and name its physical surface.

  3. Configure Canvas Size & Safety FiltersChoose from five production-standard aspect ratios (1:1, 4:3, 3:4, 16:9, 9:16) for text-to-image tasks, and optionally adjust enable_safety_checker for compliance.

  4. Submit Request & Retrieve OutputSend an authenticated POST request with your API key to submit the task, receive a task_id, and retrieve the completed high-resolution image URL.

Pricing

Each generation costs 2 credits ($0.010). Text-to-image and image editing use the same rate.

UsageRateDetails
Standard generation2 credits / $0.010Per generation

Use Cases

  • E-Commerce Lifestyle & Product PhotographyQuickly generate diverse commercial staging, studio tabletop setups, and contextual lifestyle backgrounds for consumer goods and apparel.

  • Bilingual Advertising Posters & BannersLeverage crisp dual-language character rendering to produce promotional banners, storefront signage, and marketing creatives with readable headlines.

  • High-Frequency Concept Ideation & BrainstormingRely on sub-second rendering speeds to rapidly iterate artistic concepts, mood boards, and visual design proposals in real time.

  • Controlled Image Restyling & Secondary CreationSupply a reference photo URL to alter lighting moods, swap environments, or transform visual aesthetics while preserving the subject's key composition.

  • Multi-Platform Social Media CreativesProduce on-brand visuals tailored to 1:1 square feeds, 16:9 wide headers, or 9:16 vertical full-screen formats for major social channels.

  • High-Throughput Web & Conversational Agent PipelinesSeamlessly integrate responsive, low-cost image generation into interactive chatbots, automated content engines, and customer-facing apps.

Tips

  • Wrap any target text in quotation marks within your prompt and explicitly describe its substrate, such as 'a minimalist ceramic cup with "Vidgo" printed in gold lettering'.
  • Structure prompts logically by describing the subject first, followed by environmental composition, lighting direction, and material textures to help the 8-step model lock into the scene.
  • When performing single-image editing, supply a clear, well-lit reference image and describe what to retain alongside the modifications for controlled transformations.
  • Set your desired aspect ratio via the size parameter before submitting text-to-image prompts so the spatial perspective matches your target display layout.
  • Take advantage of the economical flat rate of 2 credits ($0.010) per run to perform rapid multi-round prompt iterations and converge on your optimal visual.

Notes

  • Prompt requirements: prompt is required, accepting 1 to 1000 characters with at least one non-whitespace character.
  • Single-image editing input: image_urls is optional and accepts an array containing exactly one public HTTP(S) image URL.
  • Aspect ratio support: size accepts 1:1, 4:3, 3:4, 16:9, and 9:16; it is required for text-to-image and optional when image_urls is provided.
  • Transparent flat pricing: Both text-to-image and image editing cost 2 credits ($0.010) per generation, with automatic refunds if a task encounters an error.

Related Models

Z-Image Turbo API frequently asked questions

What is the Z-Image Turbo API?

Z-Image Turbo is a Tongyi-MAI model for fast text-to-image generation and single-image editing. It synthesizes photorealistic visuals from text prompts or single reference images in just 8 inference steps, delivering sub-second response times, accurate bilingual typography, and versatile aspect ratios. Built on a 6-billion-parameter Single-Stream Diffusion Transformer (S3-DiT) with latent distillation, it faithfully preserves prompt details and source image composition while adding rich, natural lighting. You can call it programmatically or try it from the playground above.

How many inference steps does Z-Image Turbo require?

Z-Image Turbo requires only 8 function evaluations (8 NFEs / 8 denoising steps) to synthesize high-quality images. Unlike traditional diffusion models that iterate over dozens of steps, this distilled architecture preserves fine details and sharp contrast while delivering sub-second generation speeds.

How well does Z-Image Turbo handle bilingual English and Chinese text?

Z-Image Turbo is natively optimized for bilingual typography, rendering crisp and accurate English words and Chinese characters on storefronts, book covers, packaging, and UI mockups. For optimal results, enclose the desired text in quotation marks within your prompt and name the physical object it appears on.

What aspect ratios does Z-Image Turbo support?

Z-Image Turbo natively supports five common aspect ratios: 1:1 (square), 4:3 (standard landscape), 3:4 (standard portrait), 16:9 (widescreen landscape), and 9:16 (vertical mobile). In text-to-image requests, supply the desired option in the size parameter to match your target layout.

How does the single-image editing mode work in Z-Image Turbo?

To activate single-image editing, include an array with exactly one public HTTP(S) image URL in the image_urls field. The model preserves the base composition and subject outlines of your source image while applying adjustments, style transfers, or lighting changes directed by your prompt, with no size parameter required.

How do I configure the safety checker in Z-Image Turbo?

Z-Image Turbo includes an optional enable_safety_checker boolean parameter that defaults to true to filter potentially sensitive or policy-violating imagery. If your enterprise pipeline or internal compliance framework manages safety independently, you can explicitly configure this boolean flag on submission.

Can I use images generated by Z-Image Turbo commercially?

Yes, commercial use is permitted. Z-Image Turbo is released under the open Apache 2.0 license, and generated visual assets can be used for commercial advertising, marketing campaigns, client deliverables, and production software applications.

How does Z-Image Turbo perform on photorealistic portraits and lighting?

Z-Image Turbo excels at photorealistic portraiture, capturing subtle skin pores, facial expressions, realistic hair strands, and balanced ambient lighting. Describing lighting style, camera angle, and depth of field in your prompt yields results on par with dedicated commercial photography shoots.