Wan 2.7 Text to Video API

alibaba/wan-2.7/text-to-video

Wan 2.7 Text to Video turns written prompts into 5–15 second videos at 720p or 1080p, featuring Thinking Mode prompt reasoning, native audio synchronization, and natural physical motion. It maintains scene coherence and cinematic camera movement across multiple aspect ratios while following detailed narrative instructions.

Input
446/5,000
Audio Url (optional)
Output
Ready
720p · 5 s × 12 credits/s = 60 credits ($0.300)
Continue using

Examples

One continuous five-second low water-level camera push through a large natural sea arch. A single turquoise swell enters a shadowed basalt sea cave and curls into white foam around the rocks while sunlight glows through the arch ahead. Keep the arch geometry fixed and the water physically coherent. Gradual forward motion, no cuts. Natural surf ambience. No text, lettering, captions, logos, watermarks, brands, advertising or product packaging.

An adult cellist in simple dark clothing plays one slow sustained bow stroke in a quiet sunlit rehearsal room. Five-second continuous medium side shot drifting gently sideways. Preserve natural hands, bow contact with strings, instrument geometry and the seated posture. A single warm cello note and soft room ambience. No cut. No text, lettering, captions, logos, watermarks, brands, advertising or product packaging.

Inside a mountaintop astronomical observatory at night, the curved roof shutter slowly slides open to reveal a crisp star-filled sky. Five-second continuous upward-looking wide shot, gentle camera tilt follows the opening. Keep the telescope and dome mechanically consistent, subtle cool starlight enters. Soft motor hum. No time lapse, no cuts. No text, lettering, captions, logos, watermarks, brands, advertising or product packaging.

Wan 2.7 Text to Video Overview

Wan 2.7 Text to Video translates natural language instructions into high-fidelity video sequences up to 15 seconds. Built on Alibaba Tongyi Lab's Diffusion Transformer architecture with Wan-VAE temporal compression, it analyzes complex descriptive prompts through Thinking Mode to render coherent scene physics, lighting transitions, and precise camera choreography without requiring input media.

Why Choose Wan 2.7 Text to Video?

  • Generate from text prompts directlyCompose subjects, lighting, perspective, and motion dynamics entirely through text without preparing prior visual assets.

  • Leverage Thinking Mode reasoningThe integrated reasoning engine interprets complex multi-sentence scene descriptions and preserves spatial-temporal consistency.

  • Select 5, 10, or 15-second durationsChoose the exact clip duration required for your production timeline with consistent pacing and no frame stutter.

  • Adapt to five native aspect ratiosOutput in widescreen 16:9, social 9:16, square 1:1, or classic 4:3 and 3:4 formats to match target viewing channels.

  • Connect synchronized audio tracksProvide an optional audio URL to coordinate background music or sound effects with visual beats and camera cuts.

Parameters

ParameterRequirementDescription
promptYes

Prompt, trimmed; maximum 5,000 Unicode characters. Optional for image-to-video.

Default
audio_urlNo

Optional public HTTP(S) audio URL.

Default
aspect_ratioNo

16:9, 9:16, 1:1, 4:3 or 3:4. Default: 16:9.

Default16:9
resolutionNo

720p or 1080p. Default: 720p.

Default720p
durationNo

5, 10 or 15 seconds; default 5.

Default5
seedNo

Optional integer from 0 to 2147483647.

Default
enable_safety_checkerNo

Optional safety checker setting; omitted by default.

Default

How to Use Wan 2.7 Text to Video

  1. Write a detailed scene promptDescribe the primary subject, camera movement, environment lighting, and temporal sequence clearly in the prompt.

  2. Configure duration and formatChoose the duration (5, 10, or 15 seconds), resolution (720p or 1080p), and aspect ratio matching your target distribution channel.

  3. Submit task and fetch outputExecute the generation request via API or Playground, monitor the asynchronous task status, and retrieve your final MP4 video.

Pricing

Credits = output seconds × resolution rate. Failed tasks are refunded automatically.

UsageRateDetails
720p12 credits/s ($0.06/s)All four workflows use the same rate.
1080p18 credits/s ($0.09/s)All four workflows use the same rate.

Best Use Cases

  • Cinematic pre-visualizationPrototype lighting, camera blocking, and scene transitions directly from screenplay excerpts before physical production.

  • Commercial concept teasersGenerate visual mood boards and rapid campaign proofs-of-concept for advertising pitches and client presentations.

  • Social media content creationProduce punchy 9:16 vertical video narratives and attention-grabbing social media visuals with crisp 1080p clarity.

  • Creative storytelling & fictionBring fantasy, sci-fi, and historical narrative sequences to life from descriptive writing without 3D animation pipelines.

Pro Tips for Wan 2.7 Text to Video

  • Separate subject and camera instructions:Structure your prompt with subject action in the opening clause, followed by camera angle, movement speed, and lighting tone.
  • Use temporal cues for pacing:Specify pacing markers such as slow-motion, smooth pan, or gradual zoom to guide the Diffusion Transformer backbone smoothly.
  • Match resolution to output requirements:Use 720p for rapid iterative prototyping and switch to 1080p for final broadcast-quality rendering to balance credit costs.
  • Fix the seed for creative variations:Keep the seed number constant while modifying specific adjectives or camera verbs to observe controlled composition changes.

Notes

  • Prompt character capacity:Prompts support up to 5,000 Unicode characters. Ensure descriptions are focused on visual elements rather than abstract concepts.
  • Duration selection:Text to Video supports 5, 10, or 15 seconds. Requests with other values will fail schema validation.
  • Automatic credit refund:If a task fails during generation or media processing, all deducted credits are refunded immediately to your account.

Wan 2.7 Text to Video API frequently asked questions

What is the Wan 2.7 Text to Video API?

Wan 2.7 Text to Video is an Alibaba Tongyi Lab model for text-to-video generation. It creates 5 to 15-second high-definition videos at 720p or 1080p directly from text prompts, featuring Thinking Mode prompt reasoning, native audio coordination, and natural physical motion. Built on a Diffusion Transformer backbone with Wan-VAE 3D causal temporal compression, it preserves scene coherence and cinematic camera movement while following detailed narrative prompts. You can call it programmatically or try it from the playground above.

How long can Wan 2.7 Text to Video generate in a single request?

Wan 2.7 Text to Video supports three discrete duration options: 5 seconds, 10 seconds, and 15 seconds. You can select the duration tier directly in your request payload to balance narrative pacing and generation credits.

Does Wan 2.7 Text to Video support synchronized audio generation?

Yes, Wan 2.7 Text to Video supports audio integration by accepting an optional audio_url parameter. When provided with a public audio track, the model aligns visual pacing and cinematic beats with the supplied sound.

What resolutions does Wan 2.7 Text to Video offer?

Wan 2.7 Text to Video provides 720p and 1080p native video resolutions. The default tier is 720p at 12 credits per second, while 1080p delivers crisper textures and enhanced sharpness at 18 credits per second.

Which aspect ratios does Wan 2.7 Text to Video support?

Wan 2.7 Text to Video supports five native aspect ratios: 16:9 widescreen, 9:16 vertical, 1:1 square, 4:3 standard, and 3:4 portrait. The canvas composition is generated natively according to the selected ratio without letterboxing.

How does Thinking Mode improve Wan 2.7 Text to Video generation?

Thinking Mode allows Wan 2.7 to deeply analyze the semantic hierarchy of prompts before generation. It resolves spatial layout, object relationships, and sequential actions over time to produce accurate physical interactions and smooth transitions.

How does Wan 2.7 Text to Video billing work for failed tasks?

Generation credits are charged based on the requested output seconds and resolution rate. If a task fails validation or encounters an execution error during processing, all deducted credits are immediately and fully refunded to your account.