Wan 2.6 Text to Video API

alibaba/wan-2.6/text-to-video

Wan 2.6 Text to Video transforms natural-language prompts into 5–15 second 1080p high-fidelity videos, featuring multi-shot narrative composition, cinematic camera movement, and expressive scene lighting. It faithfully preserves narrative pacing and real-world physical dynamics while sustaining temporal continuity across complex multi-angle transitions.

Input
638/5000
OutputReady
720p · 10 seconds · 160 credits = $0.800

Examples

A realistic indoor bouldering gym in bright diffuse daylight. One adult climber wearing a plain rust-red T-shirt, charcoal trousers and climbing shoes, close to the ground on a gently angled wall. In one continuous five-second side-tracking shot, the climber shifts weight onto the left foot, reaches the right hand to the next blue hold and moves one foothold sideways. Controlled modest movement, correct limb anatomy, convincing contact and body weight. Chalky wall texture, soft shoe scuff and breathing. No jumps, cuts, text, logos or music.

A realistic community kitchen in warm morning daylight. A cook in a plain pale-green apron lifts the lid of a large round bamboo steamer on a stainless counter with both hands, revealing six white steamed buns in one neat layer. One steady medium shot with a gentle push-in. Steam rises and thins naturally, the heavy lid moves as one rigid object, hands keep a secure grip. Soft lid clack, faint kitchen room tone, no dialogue or music. An ordinary shared breakfast preparation, not a product shot. No text, logos, cuts or sales imagery.

Wan 2.6 Text to Video

Wan 2.6 Text to Video is developed by Alibaba Tongyi Lab to generate cinematic dynamic video directly from natural-language text prompts. With native support for 5, 10, or 15-second takes, 1080p high-definition output, and optional multi-shot camera sequencing, it enables creators to produce complete narrative clips without requiring initial source imagery.

Why Choose This?

  • Prompt-driven multi-shot storytellingStructure story beats with natural language and enable multi_shots for automated multi-angle scene progression in a single run.

  • Up to 1080p native full HDChoose between 720p and 1080p resolutions to render subtle facial expressions, fabric textures, and dynamic environment lighting.

  • Flexible 5–15 second durationsGenerate 5, 10, or 15-second clips to accommodate everything from quick concept boards to sustained narrative takes.

  • Coherent physical motion simulationBuilt on Diffusion Transformer architecture to simulate realistic fluid flow, gravity, cloth movement, and character kinetics.

  • Seamless API integration and playgroundTest prompts in the web playground or deploy via REST API with asynchronous job submission, polling, and webhook support.

Parameters

ParameterRequirementDescription
promptRequired

String. Trimmed before validation; 1–5,000 Unicode characters.

durationOptional

Integer. Output duration: 5, 10, 15 seconds.

Default51015
resolutionOptional

Output resolution: 720p or 1080p.

Default720p1080p
multi_shotsOptional

Optional boolean. Omit to leave the upstream setting unspecified. The playground starts with false. No additional charge.

truefalse

How to Use

  1. Draft your text promptDescribe the subject, environment, action sequence, and camera movement in prompt, supporting 1–5,000 characters.

  2. Select output durationChoose a duration of 5, 10, or 15 seconds depending on your scene requirements, with 5 seconds as default.

  3. Choose resolution tierPick 720p for fast exploration or 1080p for final production delivery, with 720p as default.

  4. Optionally enable multi-shotSet multi_shots to true to introduce multi-angle cinematography without any additional cost.

  5. Verify credits and submitCheck the estimated cost and click Run, or send a POST request to /api/generate/submit.

  6. Track task progressPoll the status endpoint with task_id until status reaches finished to retrieve your download URL.

  7. Preview and downloadWatch the generated video directly in the playground player and download the MP4 file.

Pricing

Billed per generated video by output resolution and duration. Multiple shots do not add a surcharge. 1 credit = $0.005.

UsageRateDetails
720p · 5 seconds80 credits / $0.40Per video
720p · 10 seconds160 credits / $0.80Per video
720p · 15 seconds240 credits / $1.20Per video
1080p · 5 seconds120 credits / $0.60Per video
1080p · 10 seconds240 credits / $1.20Per video
1080p · 15 seconds360 credits / $1.80Per video

Best Use Cases

  • Cinematic storyboards and previsualizationTurn screenplays into dynamic previs reels with multi-shot staging to test blocking and timing.

  • Digital advertising and product promosProduce high-impact 1080p commercials and social clips from narrative and lifestyle descriptions.

  • Short-form drama and narrative contentDevelop scripted dialogue moments and sequential dramatic action across sustained 15-second takes.

  • Natural landscapes and atmosphere reelsDepict atmospheric changes such as sunrise, ocean swell, or mountain mists with organic light physics.

  • Artistic VFX and visual concept explorationExplore surreal concepts and fantastical creatures with stable physical motion and lighting.

Pro Tips

  • Structure prompts by layers: Organize descriptions into subject details, environment, action chronology, camera motion, and lighting style for optimal model comprehension.
  • Leverage multi_shots for scene progression: When describing sequence changes like close-up to wide shot, set multi_shots=true to trigger natural multi-angle framing.
  • Align duration with scene scope: Use 5 seconds for a single punchy beat, and choose 10 or 15 seconds when characters perform multi-phase actions.
  • Use standard cinematography terms: Direct the lens with precise cues like slow push-in, orbital tracking, or low-angle pedestal for predictable camera movement.
  • Specify lighting and atmosphere: Mention soft backlight, morning mist rays, or diffused ambient bounce to bring out photorealistic 1080p textures.

Notes

  • Prompt input specifications: This endpoint is text-only; prompt accepts 1–5,000 Unicode characters after trimming surrounding whitespace.
  • Duration and resolution tiers: Supports 5, 10, and 15 seconds at 720p (default) or 1080p, billed according to selected duration and resolution.
  • No surcharge for multi-shot mode: The multi_shots parameter is an optional boolean; enabling it does not add any credit cost.
  • Asynchronous task handling: Save the returned task_id to poll generation progress or configure callback_url for automated webhook delivery.

Wan 2.6 Text to Video API Frequently Asked Questions

What is the Wan 2.6 Text to Video API?

Wan 2.6 Text to Video is an Alibaba model for text-to-video generation. It transforms natural-language text prompts into 5–15 second, up to 1080p high-definition dynamic videos with support for cinematic camera movement and multi-shot narrative composition. Built on an advanced Diffusion Transformer spatio-temporal architecture, it faithfully preserves prompt physics and narrative pacing while sustaining scene and subject coherence across angle changes. You can call it programmatically or try it from the playground above.

Can Wan 2.6 Text to Video generate 15-second videos?

Yes. The model supports duration tiers of 5, 10, and 15 seconds, with 5 seconds selected by default. Across a continuous 15-second sequence, it maintains stable lighting and subject morphology while executing multi-stage actions.

Does the multi-shot feature in Wan 2.6 Text to Video cost extra?

No, it carries no extra charge. The multi_shots parameter is an optional boolean; when enabled, the model automatically renders multi-angle cinematography following prompt cues, billed strictly by standard duration and resolution rates.

How do I generate 1080p video with Wan 2.6 Text to Video?

Simply specify 1080p in the resolution parameter. The 1080p tier renders crisp facial micro-expressions, hair textures, and diffuse lighting gradients, making it ideal for high-definition commercial deliverables.

Does Wan 2.6 Text to Video follow camera motion prompts?

Yes. Adding standard camera directions such as slow push-in, orbit shot, or low-angle tracking to your prompt guides the model to execute smooth, cinematic camera paths.

What is the prompt character limit for Wan 2.6 Text to Video?

The prompt parameter accepts 1 to 5,000 Unicode characters after trimming whitespace. It natively supports English and Chinese, offering ample space for detailed scene choreography and visual direction.

Should I use Wan 2.6 Text to Video if I already have reference images?

If you have existing artwork or photo assets, choose the Wan 2.6 Image to Video endpoint instead. Wan 2.6 Text to Video is optimized for text-only creation, whereas Image to Video anchors the initial frame to preserve facial likeness and framing.