Wan 2.2 Fast Text to Video API

alibaba/wan/v2.2-a14b/text-to-video/turbo

Wan 2.2 Fast Text to Video transforms natural language prompts into fluid, cinematic video clips across 480p and 720p resolutions with rapid generation speed and flexible widescreen and portrait aspect ratios. It accurately interprets complex scene descriptions and camera motion cues while maintaining physical consistency and temporal coherence throughout the sequence.

Input

530/800

Output
Example output
480p · 6 credits ($0.030) / generation
Continue with

Examples

Full-body portrait-format shot of one adult contemporary dancer in a flowing deep burgundy skirt in a quiet sunlit rehearsal studio with a wooden floor and tall windows, no mirrors. She makes one slow graceful turn on the spot, then settles into a balanced standing pose as the skirt falls naturally around her legs. Her feet remain in contact with the floor and her arms move gently with the turn. Locked camera, entire body and feet always visible, soft daylight, realistic anatomy, one continuous shot, no other people, no text or logos.

Handcrafted clay stop-motion animation in a warm miniature library. Close medium shot, chest up, of one small teal clay robot nestled between tall worn books. The robot has a round clay head, two large amber eyes, a short neck and a simple solid torso, with visible clay fingerprints. At first its face looks toward the books on the right. In one gentle continuous action it slowly turns its head toward the camera and holds a curious attentive expression. Its torso remains still. Locked camera, shallow depth of field, soft dusty window light across tactile paper and clay surfaces, the same single robot and face throughout. Cozy and charming, no scene cuts, no text or logos.

Wan 2.2 Fast Text to Video

Wan 2.2 Fast Text to Video is an ultra-fast text-to-video generation model developed by Alibaba Tongyi Lab. Designed for rapid creative iteration, programmatic social video generation, and production pipelines, it turns natural language descriptions into high-quality video clips. Creators can render in standard 16:9 widescreen or 9:16 vertical formats across 480p and 720p resolution tiers, backed by accessible per-generation pricing starting at 6 credits ($0.030).

Why choose this model

  • High-Efficiency MoE ArchitectureLeverages a 27B Mixture-of-Experts foundation activating 14B parameters per step, delivering fast inference times without sacrificing physical simulation fidelity.

  • Flexible 16:9 and 9:16 FormatsNatively supports 16:9 landscape for widescreen cinematic video and 9:16 portrait for mobile formats like TikTok, Instagram Reels, and YouTube Shorts.

  • Dual 480p and 720p Resolution TiersOffers 480p for ultra-fast, budget-friendly drafting and 720p for crisp, ready-to-publish digital media production.

  • Nuanced Prompt InterpretationProcesses up to 800 Unicode characters of scene choreography, accurately rendering camera movements, environmental physics, and character dynamics.

  • Predictable Low-Cost PricingFixed per-generation billing at 6 credits ($0.030) for 480p and 12 credits ($0.060) for 720p, with automatic credit refunds on unexpected task failures.

Parameters

ParameterRequirementDescription
promptYes

Required non-empty string, up to 800 Unicode characters after trimming surrounding whitespace.

Default-
aspect_ratioNo

Output aspect ratio: 16:9 or 9:16.

Default16:9
resolutionNo

Output resolution: 480p or 720p.

Default720p
seedNo

Optional random seed for reproducibility.

Default-

How to Use

  1. Compose Detailed Scene PromptWrite clear prompt text up to 800 characters describing the subject appearance, background environment, lighting ambiance, and specific camera motion.

  2. Choose Aspect RatioSelect 16:9 for cinematic landscape screens or 9:16 for vertical mobile video feeds and social stories.

  3. Select Output ResolutionChoose 480p (6 credits) for rapid preview iterations or 720p (12 credits) for crisp production assets.

  4. Set Seed for ReproducibilityOptionally input an integer seed to replicate specific scene compositions or lighting styles across multiple generation runs.

  5. Dispatch and Retrieve VideoSubmit the generation request, monitor task status via polling or webhook callback, and fetch the final MP4 video URL once completed.

Pricing

Charged per generation. 1 credit = $0.005.

UsageRateDetails
480p6 credits/generation$0.030/generation
720p12 credits/generation$0.060/generation

Best Use Cases

  • Social Media & Short-Form ContentQuickly produce engaging vertical 9:16 video reels and TikTok clips from trend descriptions and catchy script hooks.

  • Commercial Ad Concept PrototypingTransform creative advertising concepts into dynamic visual mood boards to evaluate pacing and camera angles before costly production.

  • Digital Marketing & BannersCreate motion-rich promotional video backgrounds, product teasers, and display campaign animations directly from marketing copy.

  • Rapid Storyboard & Visual IdeationTest diverse narrative angles, lighting moods, and scene progressions with minimal latency during creative brainstorming.

Pro Tips

  • Separate Subject Action from Camera Motion:Explicitly describe what happens in the scene and how the camera moves (e.g., 'A sports car accelerates down an empty highway, camera pans smoothly tracking from the side').
  • Describe Lighting and Material Physics:Incorporate concrete lighting cues like 'golden hour rim light', 'neon reflections on wet asphalt', or 'wind blowing through tall grass' to enhance realism.
  • Validate Concept at 480p First:Run preliminary prompt experiments on the 480p tier (6 credits) before generating your final deliverable at 720p (12 credits).
  • Avoid Contradictory Motion Instructions:Keep dynamic directions straightforward and chronologically sequential within the 800-character budget for optimal temporal coherence.
  • Anchor Variations with Consistent Seeds:When refining prompt phrasing, lock the seed value to observe how individual descriptive adjustments impact the visual output.

Notes

  • Pure Text Input Interface:Wan 2.2 Fast Text to Video generates videos strictly from the prompt parameter; image inputs and audio uploads are not accepted on this endpoint.
  • Strict 800 Unicode Character Limit:The prompt field rejects inputs longer than 800 characters after whitespace trimming with HTTP 400; ensure prompt text stays within this boundary.
  • Supported Aspect Ratios and Resolutions:Aspect ratio must be '16:9' or '9:16'; resolution must be '480p' or '720p'. Custom arbitrary dimensions are not permitted.
  • Asynchronous Processing & Refund Safeguard:Generations run asynchronously via task_id polling. If a task terminates in a failed state due to system error, deducted credits are automatically refunded.

Related Models

Wan 2.2 Fast Text to Video API frequently asked questions

What is the Wan 2.2 Fast Text to Video API?

Wan 2.2 Fast Text to Video is an Alibaba model for text-to-video generation. It transforms natural language text prompts into fluid, cinematic video clips across 480p and 720p resolutions in both 16:9 widescreen and 9:16 portrait aspect ratios with rapid generation speed. You can call it programmatically or try it from the playground above.

What output resolutions are available in Wan 2.2 Fast Text to Video?

Wan 2.2 Fast Text to Video supports two output resolutions: 480p and 720p. The 480p option is ideal for fast prototyping and high-volume generation, while 720p delivers sharper detail and richer environmental textures suitable for final digital publication.

What aspect ratios does Wan 2.2 Fast Text to Video support?

Wan 2.2 Fast Text to Video supports two aspect ratios: standard 16:9 for widescreen desktop and landscape displays, and 9:16 for vertical mobile platforms including TikTok, Instagram Reels, and YouTube Shorts. You can configure this via the aspect_ratio parameter.

What is the prompt length limit for Wan 2.2 Fast Text to Video?

The prompt parameter accepts up to 800 Unicode characters after trimming leading and trailing whitespace. Prompts exceeding 800 characters are rejected with an HTTP 400 validation error, so prioritize descriptive scene cues, lighting directions, and camera actions within this budget.

How much does Wan 2.2 Fast Text to Video cost?

Pricing is calculated per completed generation based on resolution: 480p costs 6 credits ($0.030), and 720p costs 12 credits ($0.060). Credits are deducted upon submission and refunded automatically if generation terminates in a failed state.

How can you use the seed parameter in Wan 2.2 Fast Text to Video?

Passing an optional integer in the seed parameter allows you to maintain reproducible compositions and stylistic patterns across runs with the same prompt and aspect ratio. Omitting seed produces fresh random generations on each call.

What are the key differences between Wan 2.2 Fast and other Wan models?

Wan 2.2 Fast is optimized for high generation speed and cost efficiency for 480p and 720p video generation. Standard Wan models such as Wan 2.6 or Wan 2.7 prioritize multi-shot narrative control and 1080p outputs at higher compute tiers.

How do you retrieve videos from the Wan 2.2 Fast Text to Video API?

Submit a request to POST /api/generate/submit and read the returned task_id. While the status is not_started or running, poll GET /api/generate/status/{task_id}; stop at finished or failed. On success, read the video URLs from data.files[].file_url. On failure, read data.error_message.