Wan 3.0 Text-to-Video API

alibaba/wan-3.0/text-to-video

Wan 3.0 (Text-to-Video) transforms text prompts into 480p to 1080p high-definition video, supporting 2 to 30 second continuous generation, native audiovisual synchronization, and flexible aspect ratios. It preserves lifelike facial micro-expressions and camera trajectory while creating seamless sound effects and atmospheric audio.

Input

833/20000

Output

Ready
5 sec × $0.100/sec = $0.500

Continue with

Examples

A plump raccoon wearing a tiny translucent yellow rain poncho with the hood up stands at a flooded brick alley intersection at blue hour. Puddles cover the cobblestones. Several folded paper boats float in the largest puddle like tiny vehicles. The raccoon holds one autumn maple leaf as a conductor baton and directs the paper boats: first pointing left, then sweeping right, as if managing puddle traffic. Tiny raindrops tap the poncho. Camera: one continuous five-second waist-height lateral tracking shot moving left to right, staying parallel to the raccoon, gentle handheld sway, no cuts. Synchronized audio: steady rain on puddles, raccoon chitters, one distant bicycle bell. No readable text, letters, numbers, captions, labels, logos, brands, watermarks, advertisements, posters, UI screens, or product packaging. No people.

A beige office laser printer sits on a plain desk. There are no brand marks and no display text. A sandwich on a small plate sits on the input tray. The printer tray slowly swallows the whole sandwich. After a mechanical click-clunk, a slightly out-of-focus piece of toast ejects onto the output tray. The single status light blinks once, looking proud. Camera: locked-off medium shot, then a tiny punch-in on the click. Square 1:1 framing. One continuous five-second take, no cuts. Synchronized audio: paper-roller whir, a solid clunk, toast landing softly, one smug electronic beep. No readable text, letters, numbers, captions, labels, logos, brands, watermarks, advertisements, posters, UI screens, or product packaging.

A single goldfish wears an oversized vintage brass diving helmet inside a round glass bowl filled with water and two plants. The goldfish tries very hard to hold a dramatic movie-hero close-up stare, eyes wide. A few bubbles rise past the helmet visor. Then the goldfish gives one tiny solemn nod. Camera: slow vertical push-in from a medium portrait to a close-up of the helmet visor. 9:16 framing. One continuous five-second take, no cuts. Synchronized audio: underwater bubble pops, faint glass tap on the helmet, a small water slosh. No readable text, letters, numbers, captions, labels, logos, brands, watermarks, advertisements, posters, UI screens, or product packaging. No humans.

Wan 3.0 Text-to-Video

Wan 3.0 Text-to-Video generates continuous high-definition video with native synchronized audio from text prompts alone. Specify subject action, scene progression, and camera choreography to create 2 to 30-second clips at up to 1080p resolution in a single generation.

Why Choose This?

  • Text-only generationBuild complete characters, environments, and dynamic storylines directly from language without preparing or uploading source images.

  • Native audiovisual synthesisGenerate video frames and synchronized speech, ambient room tone, and action sound effects simultaneously within a single Diffusion Transformer pipeline.

  • Continuous 30-second clipsProduce extended takes up to 30 seconds with consistent subject identity, lifelike facial expressions, and complex multi-beat choreography.

  • Cinematic camera controlDirect pans, tilts, tracking shots, orbits, and crane movements while tailoring framing with adaptive, landscape, square, or vertical aspect ratios.

  • Up to 1080p resolutionSelect from 480p, 720p, and 1080p native output tiers to preserve intricate lighting, fine surface textures, and rich atmospheric details.

Parameters

ParameterRequirementDescription
promptRequired

String. Detailed description of subjects, actions, camera choreography, lighting, and sound effects; supports 1 to 20,000 characters.

durationOptional

Integer. Output duration in whole seconds between 2 and 30; playground defaults to 5 seconds.

Default5
resolutionOptional

String. Native output resolution tier; choices include 480p, 720p (default), or 1080p.

Default720p480p1080p
aspect_ratioOptional

String. Framing ratio; supports adaptive (default), 16:9, 4:3, 1:1, 3:4, and 9:16.

Defaultadaptive16:94:31:13:49:16
audioOptional

Boolean. Determines whether to synthesize a synchronized audio track alongside the video; defaults to true at no extra cost.

Defaulttruefalse
seedOptional

Integer. Random seed between 0 and 2,147,483,647 for reproducible trajectories and framing compositions.

enable_safety_checkerOptional

Boolean. Enables automated safety filtering on prompts and generated outputs; defaults to true.

Defaulttruefalse

How to Use

  1. Define the subject and premiseOpen your prompt with subject appearance, core wardrobe, and environmental context (e.g., A detective in a dark trench coat stands on a rain-slicked cyberpunk street at night).

  2. Sequence actions chronologicallyDescribe movements across sequential beats using clear temporal progression such as first, then, and finally (e.g., He inspects an illuminated sign, then walks briskly down a narrow neon alley).

  3. Direct the camera in a separate sentenceState shot size and camera motion independently from character movement (e.g., Eye-level medium tracking shot, slowly pushing in as he walks into the alleyway).

  4. Establish atmosphere and audio cuesSpecify color palette, contrast, and audio ambience (e.g., Cool blue tones with amber reflections, accompanied by footsteps splashing in puddles and distant rumbling thunder).

  5. Select resolution and ratioChoose your desired resolution (720p or 1080p), format ratio (16:9 landscape or 9:16 portrait), and length between 2 and 30 seconds.

  6. Submit and retrieve the resultRun the generation to submit an asynchronous task, then inspect the synchronized video and audio playback in the output preview.

Pricing

Wan 3.0 Text-to-Video charges by generated output second based strictly on the selected resolution tier; enabling or disabling audio carries no extra fee (1 credit = $0.005).

UsageRateDetails
480p10 credits / output sec ($0.05 / sec)Standard definition tier. 5-second default is 50 credits ($0.25); 30-second maximum is 300 credits ($1.50).
720p (Default)20 credits / output sec ($0.10 / sec)High definition tier. 5-second default is 100 credits ($0.50); 30-second maximum is 600 credits ($3.00).
1080p40 credits / output sec ($0.20 / sec)Full high definition flagship tier. 5-second default is 200 credits ($1.00); 30-second maximum is 1,200 credits ($6.00).

Best Use Cases

  • Commercial concept filmsTransform written creative scripts into dynamic concept clips for stakeholder reviews prior to live shoots.

  • Cinematic narrative previsualizationConvert screenplay scenes into continuous 2 to 30-second motion references to evaluate pacing and camera angles.

  • Multi-format social contentProduce platform-ready 16:9, 9:16, or 1:1 clips with synchronized ambient soundscapes from a single prompt idea.

  • Concept worldbuilding studiesTurn rich descriptions of imaginary creatures, futuristic cities, or natural phenomena into immersive motion studies.

Pro Tips

  • Structure prompts into distinct layers: Organize your prompt from broad setting and subject appearance to chronological actions, camera trajectory, and audio ambience.
  • Separate character actions from camera motion: Keeping character movements in one sentence and camera directions in another ensures the model interprets trajectory accurately.
  • Pace extended generations with timestamps: For takes over 10 seconds, guide narrative progression by structuring sentences with timestamps (e.g., Seconds 0-5... followed by seconds 6-10...).
  • Describe specific sounds to trigger native audio: Mentioning explicit sound cues like sharp heel clicks, mechanical hums, or reverberant dialogue guides the audio diffusion model.
  • Iterate quickly before committing long takes: Test initial prompts with 5-second 720p generations to verify movement and composition before scaling up to 1080p and 30 seconds.

Notes

  • Text-only input mode: This endpoint accepts prompt text only; image or multimodal media files are neither required nor accepted.
  • Whole-second duration input: The duration parameter accepts whole integers between 2 and 30 seconds.
  • Asynchronous lifecycle: Task creation returns a task_id immediately; retrieve the final video asset via polling or an automated webhook callback.

Wan 3.0 Text-to-Video API — Frequently Asked Questions

What is the Wan 3.0 Text-to-Video API?

Wan 3.0 Text-to-Video is an Alibaba Tongyi Lab model for generating video from text. It creates high-definition videos up to 30 seconds at up to 1080p resolution with native synchronized audio directly from natural language prompts, supporting professional cinematographic movement and flexible aspect ratios. Built on Diffusion Transformer and Flow Matching architectures, it maintains stable facial micro-expressions and complex camera trajectories while naturally aligning movement beats with ambient acoustics. You can call it programmatically or try it from the playground above.

Can Wan 3.0 Text-to-Video generate 30-second clips in a single call?

Yes. The model supports specifying any integer duration between 2 and 30 seconds (the playground defaults to 5 seconds). Throughout an uninterrupted 30-second take, it maintains facial identity and spatial coherence while unfolding multi-stage dramatic choreography.

Is Wan 3.0 Text-to-Video audio synthesized natively alongside the video?

Yes, it is generated natively. Audio tracks are denoised alongside video latents directly within the underlying diffusion transformer rather than spliced via external post-production. With audio enabled by default, the model generates synchronized ambient soundscapes, footsteps, and physical interactions derived from your text description.

Does Wan 3.0 Text-to-Video support generating silent video?

Yes. If your workflow requires silent video footage for external dubbing or editing, set the audio parameter explicitly to false. Toggling audio off does not affect rendering speed or credit pricing.

How can I maintain facial consistency during long takes in Wan 3.0 Text-to-Video?

Structure multi-beat scene choreography with chronological markers (such as first, next, finally) and describe character appearance details separately from camera motion. Avoiding contradictory actions in a single sentence helps the model maintain stable facial geometry and bodily proportions over extended durations.

How does the adaptive aspect ratio work in Wan 3.0 Text-to-Video?

When aspect_ratio is set to adaptive (default), the model intelligently determines the most harmonious composition framing based on the scene and camera descriptors in your prompt. You can also explicitly select 16:9, 4:3, 1:1, 3:4, or 9:16 to target widescreen cinema or mobile vertical feeds directly.

Are additional parameters required to output 1080p in Wan 3.0 Text-to-Video?

No extra parameters are needed. Simply select 1080p in the resolution setting (alongside 480p and 720p). Native 1080p resolution renders fine hair strands, water reflections, and architectural textures with exceptional fidelity for production delivery.