Wan 3.0 Reference-to-Video API

alibaba/wan-3.0/reference-to-video

Wan 3.0 (Reference-to-Video) transforms text prompts alongside multimodal references including images, video clips, audio, documents, or web links into dynamic video, supporting up to 20 combined assets and 2 to 30 second continuous generation. It preserves subject identity, artistic style, and narrative continuity across scenes while synthesizing synchronized native audiovisual footage.

Input

631/20000
Total references3/21 · remaining 18
3/10 · remaining 7
Reference preview
Reference preview
Reference preview
0/5 · remaining 5
0/5 · remaining 5

Output

Ready
5 sec × $0.100/sec = $0.500

Continue with

Examples

Use Reference Image 1 only for the proud pigeon with the aviator scarf. Use Reference Image 2 only for the oversized croissant. Use Reference Image 3 only for the subway grab handle. In one continuous five-second shot, the Image 1 pigeon holds the Image 2 croissant in its beak and stands on the Image 3 subway handle, swaying with the train. Camera: waist-height handheld sway matching the subway motion. No cuts. Synchronized audio: subway rumble, a proud coo, a faint pastry flake crunch. No readable text, letters, numbers, captions, labels, logos, brands, watermarks, advertisements, posters, UI screens, or product packaging.

Use Reference Image 1 only for the long-legged Victorian teapot character. Use Reference Image 2 only for the mossy library aisle. In one continuous five-second shot, the Image 1 teapot tiptoes down the Image 2 aisle, pausing to inspect a shelf, lid tilting like a polite nod. Camera: tracking follow from behind and slightly beside, staying at teapot-eye height. No cuts. Synchronized audio: porcelain foot taps, a floorboard creak, a hushed library room tone. No readable text, letters, numbers, captions, labels, logos, brands, watermarks, advertisements, posters, UI screens, or product packaging.

Parse Reference Document 1 as a one-page playful science comic about why toast usually lands butter-side down. Produce one continuous five-second explainer animation: the toast slides off a table, rotates in the air, and lands butter-side down while a tiny curious cat watches. Keep the homemade paper-comic look. No product brands, no advertisements, no logos. Camera: simple locked comic-panel push that follows the fall. No cuts. Synchronized audio: plate scrape, brief whoosh, buttered toast slap, a tiny cat chirp.

Wan 3.0 Reference-to-Video

Wan 3.0 Reference-to-Video generates continuous high-definition video with native synchronized audio by blending text prompts with multimodal reference assets. Maintain strict facial identity, costume styling, dynamic motion patterns, or document narrative structures across 2 to 30-second generations at up to 1080p resolution.

Why Choose This?

  • Omni-asset multimodal inputsCombine up to 20 multimodal assets in a single request, including images (up to 10), videos (up to 5), audio files (up to 5), and a document or webpage.

  • Cross-scene character consistencyAnchor protagonist facial features, wardrobe details, and unique aesthetic rendering styles steadily across distinct shots and story beats.

  • Document and webpage synthesisIngest structured presentations (PPT, PDF, DOCX) or public URLs to automatically distill key information into dynamic narrative video.

  • Native audiovisual co-generationSynthesize synchronized dialogue timing, acoustic room ambience, and foley sound effects natively alongside visual diffusion frames at 30 fps.

  • Extended 30-second 1080p outputDeliver broadcast-ready takes with flexible durations from 2 to 30 seconds at 480p, 720p, or 1080p resolution.

Parameters

ParameterRequirementDescription
promptRequired

String. Describes scene actions, camera motion, and assigns roles to provided reference assets; supports 1 to 20,000 characters.

reference_image_urlsOptional

Array of strings. Up to 10 image URLs for character appearance, scene environment, or props; max 30MB per image in JPEG, PNG, or WebP.

reference_video_urlsOptional

Up to 5 public HTTP(S) reference video URLs to guide motion, camera behavior, or pacing.

reference_audio_urlsOptional

Array of strings. Up to 5 audio URLs for vocal timbre, rhythm, or ambient sound; max 50MB per file in MP3 or WAV format.

reference_file_urlsOptional

Up to 1 public HTTP(S) document URL; cannot be combined with reference_link_urls.

reference_link_urlsOptional

Array of strings. Maximum 1 publicly accessible webpage URL; mutually exclusive with reference_file_urls.

durationOptional

Integer. Output duration in whole seconds between 2 and 30; playground defaults to 5 seconds.

Default5
resolutionOptional

String. Native output resolution tier; choices include 480p, 720p (default), or 1080p.

Default720p480p1080p
aspect_ratioOptional

String. Framing ratio; supports adaptive (follows reference framing), 16:9, 4:3, 1:1, 3:4, and 9:16.

Defaultadaptive16:94:31:13:49:16
audioOptional

Boolean. Determines whether to synthesize a synchronized audio track alongside the video; defaults to true at no extra cost.

Defaulttruefalse
seedOptional

Integer. Random seed between 0 and 2,147,483,647 for reproducible trajectories and dynamics.

enable_safety_checkerOptional

Boolean. Enables automated safety filtering on prompts and generated outputs; defaults to true.

Defaulttruefalse

How to Use

  1. Gather and attach reference assetsUpload character turnarounds (up to 10 images), action motion clips (up to 5 videos), timbre audio clips, or attach a pitch presentation / public webpage URL.

  2. Assign asset roles in your promptReference assets explicitly in natural language (e.g., Using the hero in reference image 1 wearing the armor in reference image 2, perform a slow walk inside the futuristic hangar from reference image 3).

  3. Choreograph motion and camera pathsDescribe scene progression sequentially with explicit camera instructions (e.g., The camera tracks steadily at waist level before ascending into an aerial view).

  4. Configure duration and resolutionSelect an output duration between 2 and 30 seconds, choose 720p or 1080p resolution, and set aspect ratio to adaptive or 16:9.

  5. Verify asset exclusivity constraintsEnsure audio is enabled if sound is desired, and verify that document files and webpage URLs are not submitted in the same request.

  6. Submit and evaluate consistent videoSubmit your task, track asynchronous progress in the preview console, and inspect character consistency upon playback before downloading.

Pricing

Wan 3.0 Reference-to-Video charges by generated output second based strictly on the selected resolution tier; multimodal asset ingestion and audio synthesis carry no additional fee (1 credit = $0.005).

UsageRateDetails
480p10 credits / output sec ($0.05 / sec)Standard definition tier. 5-second default is 50 credits ($0.25); 30-second maximum is 300 credits ($1.50).
720p (Default)20 credits / output sec ($0.10 / sec)High definition tier. 5-second default is 100 credits ($0.50); 30-second maximum is 600 credits ($3.00).
1080p40 credits / output sec ($0.20 / sec)Full high definition flagship tier. 5-second default is 200 credits ($1.00); 30-second maximum is 1,200 credits ($6.00).

Best Use Cases

  • Episodic IP character storytellingGenerate multi-shot scenes across varied environments while anchoring protagonist appearance and costume design.

  • Presentation and pitch deck visualizationConvert business proposals, pitch decks, or training slides into dynamic motion explainers with ambient voiceover soundscapes.

  • Webpage and article reformattingProvide a public article or product URL to distill written editorial content into engaging short-form social video.

  • Motion choreography style transferCombine existing stunt or dance video clips with new character portraits to transfer complex choreography onto new subjects.

Pro Tips

  • Provide multi-angle character references: Supplying front, three-quarter, and side portraits enhances subject geometric stability during dynamic character turns.
  • Use explicit reference binding phrases: Write phrases like 'character from reference image 1 wearing jacket from reference image 2' to prevent visual feature leakage.
  • Respect document and link mutual exclusivity: reference_file_urls and reference_link_urls cannot be used simultaneously; select one modality per task.
  • Keep documents focused, with clear charts and readable text hierarchy to guide the video narrative.
  • Maintain shared asset sets across shots: Reuse the exact same reference array across consecutive prompt generations to maintain consistent production design across an entire video project.

Notes

  • Reference asset limits: A single request supports up to 20 total assets: maximum 10 images, 5 videos, 5 audio files, and 1 document or webpage.
  • Mutually exclusive inputs: reference_file_urls and reference_link_urls are mutually exclusive; submitting both triggers a 400 validation error.
  • Whole-second duration input: The duration parameter accepts whole integers between 2 and 30 seconds.

Wan 3.0 Reference-to-Video API — Frequently Asked Questions

What is the Wan 3.0 Reference-to-Video API?

Wan 3.0 Reference-to-Video is an Alibaba Tongyi Lab model for generating video from multimodal references. It combines natural-language prompts with images, video clips, audio tracks, structured documents, or public webpages (up to 20 reference assets) to generate continuous takes up to 30 seconds at up to 1080p with native audio. Built on cross-modal alignment and Diffusion Transformer architectures, it preserves character facial identity, movement pacing, or document knowledge across shots in brand-new narrative scenes. You can call it programmatically or try it from the playground above.

Can I upload PPT or PDF documents to Wan 3.0 Reference-to-Video?

Yes. By providing a document URL in reference_file_urls (supporting PPT, PPTX, PDF, DOCX, TXT, MD), the model analyzes slide layout, graphics, and textual hierarchy, transforming written assets into narrated video sequences with dynamic camera moves.

Can I provide both a document and a webpage link in Wan 3.0 Reference-to-Video?

No. The reference_file_urls and reference_link_urls parameters are mutually exclusive; you may supply at most one document or one webpage link per request. You can freely combine either option with reference images, videos, or audio tracks.

What is the maximum number of reference assets in Wan 3.0 Reference-to-Video?

A single request supports up to 20 total assets, including up to 10 reference images, up to 5 reference videos, up to 5 reference audio tracks, and 1 document or webpage link. This rich multi-asset ingestion allows complex storyboard setups.

How do I bind multiple subjects in Wan 3.0 Reference-to-Video prompts?

Use explicit role-assignment tagging in your prompt, such as "Use Reference Image 1 for the protagonist's face and jacket, Reference Image 2 for the handheld prop, and Reference Video 1 for the walking pace." Explicit naming ensures distinct assets bind to the intended entities.

How does Wan 3.0 Reference-to-Video process reference audio?

When reference_audio_urls are provided, the model incorporates the reference vocal timbre, cadence, or melodic rhythm, harmonizing it with the subject's on-screen movements and ambient room acoustics for cohesive audiovisual pacing.

How do I provide a document reference?

Provide one publicly accessible HTTP(S) document URL in reference_file_urls. Do not also provide reference_link_urls.