Wan 3.0 Prime Text-to-Video API

alibaba/wan-3.0-prime/text-to-video

Wan 3.0 Prime (Text-to-Video) transforms a text prompt into a continuous video up to 30 seconds, with native audio-visual sync, layered control for action and camera, and output up to 1080p. It follows ordered beats across the shot so subject motion, scene continuity, and timing stay readable from the written plan.

Input

537/20000

Output

Idle

Your generated video will appear here

Configure the required inputs, resolution, and duration, then run the task.

5 sec × $0.140/sec = $0.700

Continue with

Examples

Vertical 9:16, 4 seconds. A fleet of bright yellow rubber ducks paddles through a canyon of stacked rolling office chairs. Low tracking camera follows the lead duck along the chair-canyon floor. Overhead fluorescent office light, no people. Native audio: tiny plastic hulls bumping, chair-wheel squeaks, shallow water splash. No logos, no readable text, no products, no packaging, no prices, no CTA, no advertising, no watermark, no brand marks.

One continuous 4-second 4:3 shot at a rainy bus shelter. Two city pigeons in transparent plastic raincoats slap each other with their wings over a soggy bread crust. Locked medium shot, rain streaking the plexiglass. Native audio: wing slaps, pigeon coos, rain on the shelter roof, a distant bus hiss. No logos, no readable text, no products, no packaging, no prices, no CTA, no advertising, no watermark, no brand marks.

Wan 3.0 Prime Text-to-Video

Wan 3.0 Prime Text-to-Video generates continuous clips with optional synchronized sound from a text prompt alone. It develops subject, setting, action order, camera language, lighting, and audio cues written in one plan so the scene advances with clear temporal continuity.

Why Choose This?

  • Text-to-VideoGenerate a continuous video from a written scene without uploading start frames or reference packs.

  • Temporal continuitySequence opening, progression, and closing beats so subject motion and scene relationships stay coherent across clips up to 30 seconds.

  • Layered prompt directionDescribe subject performance first, then control shot size, camera path, lighting, and atmosphere as separate layers.

  • Native audio syncKeep audio enabled to generate synchronized dialogue, ambience, effects, or music with the picture.

  • Aspect ratio controlChoose adaptive, 16:9, 4:3, 1:1, 3:4, or 9:16 before locking composition and negative space.

  • Delivery specsOutput 480p, 720p, or 1080p video from 2–30 seconds for draft checks and final delivery.

Parameters

ParameterRequirementDescription
promptRequired

String. Defines the scene, action, camera, visual treatment, and sound intent; 1–20,000 characters after trimming.

durationOptional

Integer. Sets output length from 2 through 30 seconds, inclusive; default is 5.

resolutionOptional

String. Sets output resolution; default is 720p.

480p720p1080p
aspect_ratioOptional

String. Controls output framing; default is adaptive.

adaptive16:94:31:13:49:16
audioOptional

Boolean. Requests a generated audio track; default is true. The audio toggle does not change the credit rate.

truefalse
seedOptional

Integer. Optional reproducibility seed from 0 through 2147483647.

enable_safety_checkerOptional

Boolean. Enables the safety checker; default is true.

truefalse

How to Use

  1. Write the scene premiseOpen with the subject, setting, and visual treatment: a potter centers clay on a wheel under soft window light.

  2. Stage ordered actionWrite concrete beats in time order: she wets the clay, lifts a tall vessel, then steadies the rim as the wheel slows.

  3. Direct camera and soundAdd shot size, angle, and movement in a separate sentence, then note dialogue, ambience, or music cues.

  4. Set durationChoose an integer from 2 through 30 seconds; the default is 5 for first drafts.

  5. Choose resolutionSelect 480p for quick motion checks, or 720p / 1080p for review and delivery.

  6. Choose aspect ratioPick adaptive, 16:9, 4:3, 1:1, 3:4, or 9:16 to match the channel framing.

  7. Configure audioKeep audio enabled for synchronized sound; turn it off when you need a silent clip.

  8. Generate the videoClick Run, then preview picture and sound together in the output area when the task finishes.

Pricing

Price depends only on output duration and resolution; the audio toggle does not change the rate.

UsageRateDetails
480p13.6 credits/output sec ($0.068/sec)2 seconds costs 27.2 credits ($0.136), 5 seconds costs 68 credits ($0.340), and 30 seconds costs 408 credits ($2.04).
720p28 credits/output sec ($0.14/sec)2 seconds costs 56 credits ($0.280), 5 seconds costs 140 credits ($0.70), and 30 seconds costs 840 credits ($4.20).
1080p56 credits/output sec ($0.28/sec)2 seconds costs 112 credits ($0.560), 5 seconds costs 280 credits ($1.40), and 30 seconds costs 1,680 credits ($8.40).

Best Use Cases

  • Campaign concept filmsTurn a written product scenario and camera plan into a concept clip for creative review before a shoot.

  • Storyboard previsualizationConvert a scripted beat into a motion reference for checking pacing, staging, and shot direction.

  • E-commerce launch teasersGenerate product storytelling clips from a launch brief for detail-page and social review.

  • Social channel variationsDevelop one written campaign idea into 9:16, 1:1, or 16:9 concepts for channel-specific review.

  • Atmosphere and music studiesTranslate lighting direction, visual progression, and sound cues into a short mood film.

Pro Tips

  • Structure the prompt as duration and aspect intent, subject and assets, scene and lighting, camera and shot, dialogue and sound, then timeline.
  • Replace a thin prompt such as 'a boutique opens' with a visible beat: staff unlock the door, warm lights rise, and the camera tracks past the display table.
  • Keep subject movement and camera movement in separate sentences so each instruction has a clear role.
  • Use first, then, and finally when several actions need a readable sequence inside a longer take.
  • Validate motion and timing at 480p / 5 seconds, then render 1080p and longer durations once the plan holds.

Notes

  • Generation is asynchronous; retain task_id and stop tracking when the task reaches finished or failed.

Wan 3.0 Prime Text To Video API — Frequently asked questions

What is the Wan 3.0 Prime Text-to-Video API?

Wan 3.0 Prime Text-to-Video is an Alibaba Tongyi Lab model for generating high-definition video from text. It creates high-fidelity videos up to 30 seconds at up to 1080p resolution with native synchronized audio from complex multi-layered prompts, featuring enhanced dynamic physical realism and precise cinematographic movement. Built on Wan 3.0 Prime's upgraded spatiotemporal reasoning architecture, it delivers exceptional character structural stability and natural camera fluidity across long cinematic takes. You can call it programmatically or try it from the playground above.

What performance enhancements does Wan 3.0 Prime Text-to-Video offer over Wan 3.0?

Wan 3.0 Prime delivers significant upgrades in character anatomy stability, fine-grained physics (such as liquid dynamics, fabric folds, and particles), and cinematographic lighting consistency over Wan 3.0. In complex multi-beat scenes, Prime follows layered action instructions more faithfully with reduced artifact distortion.

Can Wan 3.0 Prime Text-to-Video generate 30-second videos in a single task?

Yes. The model natively supports specifying any duration between 2 and 30 seconds (defaulting to 5 seconds). Across a 30-second single take, the model maintains consistent character identity, continuous motion trajectories, and natural temporal rhythm.

Is audio generated natively alongside video in Wan 3.0 Prime Text-to-Video?

Yes. Audio latents are co-synthesized within the multimodal diffusion framework alongside video frames. When audio is set to true (default), the model automatically infers environmental acoustics, Foley sounds, and action interactions from your prompt text.

How should I prompt multi-beat camera moves in Wan 3.0 Prime Text-to-Video?

We recommend organizing prompts into chronological stages using cues such as "first", "then", and "finally". Separate character choreography from camera movement instructions (e.g. low-angle tracking shot, slow push-in) to ensure the model executes each motion phase coherently.

Does Wan 3.0 Prime Text-to-Video support all major aspect ratios?

Yes. The model provides adaptive aspect ratio alongside 16:9 widescreen, 4:3 standard, 1:1 square, 3:4 portrait, and 9:16 vertical formats. Adaptive framing automatically selects the optimal composition based on scene content, while 9:16 targets vertical social feeds directly.

Can I draft prompts with Wan 3.0 Prime Text-to-Video in 480p to conserve credits?

Yes. The model offers 480p, 720p, and 1080p tiers. Generating at 480p consumes only 13.6 credits/s, making it ideal for verifying prompt choreography, scene staging, and camera pacing before committing to full 1080p production rendering.