Wan 3.0 Prime Reference-to-Video API

alibaba/wan-3.0-prime/reference-to-video

Wan 3.0 Prime (Reference-to-Video) transforms a text prompt plus image, video, audio, file, or link references into a continuous video up to 30 seconds, with Omni-Reference multimodal control, native audio-visual sync, and output up to 1080p. It carries identity, props, motion cues, and space into a new scene while following the roles you assign in the prompt.

Input

500/20000
Total references0/20 · remaining 20
0/10 · remaining 10
0/5 · remaining 5
0/5 · remaining 5

Add at least one reference image, video, audio, document URL, or webpage URL.

Output

Idle

Your generated video will appear here

Configure the required inputs, resolution, and duration, then run the task.

5 sec × $0.140/sec = $0.700

Continue with

Examples

Use the two reference images only for this long-legged porcelain teapot: spout, lid, floral glaze, and stork-like legs. New 4-second 1:1 scene: the teapot stomps across a wooden dining table, each footfall making cups rattle. Locked three-quarter camera. Native audio: ceramic foot stomps, cup chatter, table creak. Image references only. No logos, no readable text, no products, no packaging, no prices, no CTA, no advertising, no watermark, no brand marks.

Use the two close-up references for this barrel cactus with tiny black sunglasses, spines, and pot. New 4-second 4:3 scene: the cactus nods and sways in a tiny dance on a sunny windowsill. Locked close-up. Native audio: faint spine rustle, window breeze, a soft rhythmic tap. No logos, no readable text, no products, no packaging, no prices, no CTA, no advertising, no watermark, no brand marks.

Wan 3.0 Prime Reference-to-Video

Wan 3.0 Prime Reference-to-Video generates continuous clips with optional synchronized sound from a required prompt and mixed references. It brings subject appearance, motion rhythm, camera language, and sonic cues from those assets into a new scene when each reference is given a clear role.

Why Choose This?

  • Reference-to-VideoCombine a prompt with image, video, audio, file, or link references to generate a new continuous scene.

  • Subject continuityCarry a recurring character, product, voice, or room into a prompt-driven clip instead of rebuilding identity from text alone.

  • Multimodal reference rolesUse up to 10 images, 5 videos, and 5 audio clips, or one file / one link, and assign each asset a clear creative job.

  • Native audio syncKeep audio enabled to generate synchronized dialogue, ambience, effects, or music with the picture.

  • Layered prompt directionMap each reference to identity, motion, camera, or sound, then describe the new scene timeline.

  • Delivery specsOutput 480p, 720p, or 1080p video from 2–30 seconds with adaptive or fixed aspect ratios.

Parameters

ParameterRequirementDescription
promptRequired

String. Defines how each reference contributes to the new scene; 1–20,000 characters after trimming.

reference_image_urlsConditional

String array of 1–10 public image URLs. Uses images to guide identity, appearance, composition, environment, or style.

reference_video_urlsConditional

String array of 1–5 public video URLs. Uses videos to guide action, camera movement, blocking, pace, or shot rhythm.

reference_audio_urlsConditional

String array of 1–5 public audio URLs. Uses audio to guide ambience, rhythm, voice character, or sound direction.

reference_file_urlsConditional

String array with exactly one public file URL when a packaged reference file is the source. Mutually exclusive with reference_link_urls.

reference_link_urlsConditional

String array with exactly one public link URL when a linked reference pack is the source. Mutually exclusive with reference_file_urls.

durationOptional

Integer. Sets output length from 2 through 30 seconds, inclusive; default is 5.

resolutionOptional

String. Sets output resolution; default is 720p.

480p720p1080p
aspect_ratioOptional

String. Controls output framing; default is adaptive.

adaptive16:94:31:13:49:16
audioOptional

Boolean. Requests a generated audio track; default is true. The audio toggle does not change the credit rate.

truefalse
seedOptional

Integer. Optional reproducibility seed from 0 through 2147483647.

enable_safety_checkerOptional

Boolean. Enables the safety checker; default is true.

truefalse

How to Use

  1. Decide each asset's roleLabel whether a reference controls character identity, product look, environment, motion, camera path, voice, or rhythm.

  2. Add the reference anchorsProvide at least one of reference_image_urls, reference_video_urls, reference_audio_urls, reference_file_urls, or reference_link_urls within the allowed counts.

  3. Map relationships in the promptWrite the mapping plainly: first image for the presenter, first video for walking pace, first audio for room tone.

  4. Set durationChoose an integer from 2 through 30 seconds; the default is 5 for first drafts.

  5. Choose resolutionSelect 480p for role checks, or 720p / 1080p for review and delivery.

  6. Choose aspect ratioPick adaptive, 16:9, 4:3, 1:1, 3:4, or 9:16 to match the channel framing.

  7. Configure audioKeep audio enabled for synchronized sound; turn it off for a silent clip.

  8. Generate the videoClick Run, then preview which reference roles carried into the clip in the output area.

Pricing

Price depends only on output duration and resolution; the audio toggle does not change the rate.

UsageRateDetails
480p13.6 credits/output sec ($0.068/sec)2 seconds costs 27.2 credits ($0.136), 5 seconds costs 68 credits ($0.340), and 30 seconds costs 408 credits ($2.04).
720p28 credits/output sec ($0.14/sec)2 seconds costs 56 credits ($0.280), 5 seconds costs 140 credits ($0.70), and 30 seconds costs 840 credits ($4.20).
1080p56 credits/output sec ($0.28/sec)2 seconds costs 112 credits ($0.560), 5 seconds costs 280 credits ($1.40), and 30 seconds costs 1,680 credits ($8.40).

Best Use Cases

  • Brand-kit filmsCombine product stills, location photos, short motion refs, and a logo lockup into one on-brand clip driven by a new prompt.

  • Character or product continuityCarry a character sheet, wardrobe, or product hero into a new performance or demo scene.

  • Deck or report explainersUse one reference file or public webpage pack to turn structured content into a narrated explainer clip.

  • Motion and voice-matched cutsPair a subject image with a movement video and a voice or rhythm track for campaign spots that follow those references.

  • Multi-asset storyboard assemblyAssemble approved visual, motion, and audio references into a single previsualization take for creative review.

Pro Tips

  • Write the role beside each asset in the prompt: character, wardrobe, product, room, motion, camera, voice, or tempo.
  • Replace 'use all references' with a relationship: keep the subject from the first image, follow the blocking in the first video, and use the first audio only for tempo.
  • Stay inside the allowed caps: up to 10 images, 5 videos, 5 audio clips, and exactly one file or one link when using document or webpage input.
  • Structure the prompt as duration and aspect intent, subject and reference assets, scene and lighting, camera and shot, dialogue and sound, then timeline.
  • Validate reference roles at 480p / 5 seconds, then render 1080p and longer durations once identity and timing hold.

Notes

  • Provide at least one reference array; reference_file_urls and reference_link_urls are mutually exclusive.
  • Generation is asynchronous; retain task_id and stop tracking when the task reaches finished or failed.

Wan 3.0 Prime Reference To Video API — Frequently asked questions

What is the Wan 3.0 Prime Reference-to-Video API?

Wan 3.0 Prime Reference-to-Video is an Alibaba Tongyi Lab model for generating high-definition video from multimodal references. It combines natural-language prompts with multiple images, video clips, audio tracks, structured documents, or public webpages (up to 20 reference assets) to generate continuous takes up to 30 seconds at up to 1080p with native audio. Built on Wan 3.0 Prime's upgraded multimodal fusion architecture, it maintains strict character facial identity, movement rhythm, and visual style across shots in brand-new narrative environments. You can call it programmatically or try it from the playground above.

What multi-asset fusion upgrades does Wan 3.0 Prime Reference-to-Video offer over Wan 3.0?

Wan 3.0 Prime features enhanced multi-modal asset feature alignment and cross-modal reasoning. When ingesting multiple references (such as character portraits, prop photos, and motion reference clips simultaneously), it maps identity, motion dynamics, and acoustic cues with significantly higher fidelity and visual cohesion.

Can Wan 3.0 Prime Reference-to-Video ingest PPT or PDF documents?

Yes. By providing a document URL in reference_file_urls (supporting PPT, PPTX, PDF, DOCX, TXT, MD), the model analyzes slide graphics, layout hierarchy, and textual knowledge to produce structured, narrated video presentations.

Can I provide both a document and a webpage link in Wan 3.0 Prime Reference-to-Video?

No. The reference_file_urls and reference_link_urls parameters are mutually exclusive; you may submit at most one document or one webpage link per request. You can freely combine either option with reference images, videos, or audio tracks.

How do I assign roles to multiple references in Wan 3.0 Prime Reference-to-Video?

Use clear role-assignment tagging in your prompt, such as "Use Reference Image 1 for the character's face, Reference Image 2 for the costume, and Reference Video 1 for the walking pace." Explicit mapping guides the model to bind each asset to the correct entity.

How does Wan 3.0 Prime Reference-to-Video synthesize native audio?

When reference_audio_urls are provided, the model incorporates reference vocal tone, cadence, or ambient music into the scene. If no audio assets are supplied, it generates matching environmental soundscapes and Foley effects derived from prompt and visual cues.

What are the reference image specifications for Wan 3.0 Prime Reference-to-Video?

You can include up to 10 reference images (JPEG, PNG, WebP up to 30MB each). Supplying multi-angle views—such as front portraits, profile perspectives, and wardrobe details—gives the model comprehensive geometric data to maintain character consistency across dynamic takes.