Wan 3.0 Prime Image-to-Video API

alibaba/wan-3.0-prime/image-to-video

Wan 3.0 Prime (Image-to-Video) transforms a start-frame image and text prompt into a continuous video up to 30 seconds, with optional end-frame guidance, native audio-visual sync, and output up to 1080p. It carries the source subject, composition, and style into motion while adding action, camera movement, and sound.

Input

514/20000

Output

Idle

Your generated video will appear here

Configure the required inputs, resolution, and duration, then run the task.

5 sec × $0.140/sec = $0.700

Continue with

Examples

Begin from the first frame. One continuous 4-second 16:9 shot: the eyed toast slides straight toward the lens on a thin butter trail, eyes widening, until it nearly fills the frame. Locked low table-level camera. Native audio: bread scrape on laminate, a tiny squeak, quiet kitchen room tone. No logos, no readable text, no products, no packaging, no prices, no CTA, no advertising, no watermark, no brand marks.

Begin from the first frame. One continuous 4-second 3:4 shot: the diving-helmet cat slowly turns its head toward the goldfish in the open briefcase. Preserve the same cat, helmet, briefcase, fish, and desk. Locked camera. Native audio: helmet metal tick, water slosh in the pouch, soft room tone. No logos, no readable text, no products, no packaging, no prices, no CTA, no advertising, no watermark, no brand marks.

Wan 3.0 Prime Image-to-Video

Wan 3.0 Prime Image-to-Video generates continuous clips with optional synchronized sound from a start-frame image and an optional text prompt. It preserves the opening subject, composition, and lighting while adding subject motion, camera movement, scene progression, and sound; an optional second image can guide the closing state.

Why Choose This?

  • Image-to-VideoAnimate an approved still into a continuous clip by providing one or two public image URLs.

  • Preserve source featuresCarry the start frame's subject identity, composition, lighting, and style into the generated motion.

  • Optional end-frame guidanceAdd a second image in image_urls when the final pose, product state, or composition needs a visible destination.

  • Native audio syncKeep audio enabled to generate synchronized ambience, action sounds, dialogue, or music with the animated still.

  • Prompt-led motion controlDescribe subject action, camera path, lighting change, and atmosphere after the opening frame.

  • Delivery specsOutput 480p, 720p, or 1080p video from 2–30 seconds with adaptive or fixed aspect ratios.

Parameters

ParameterRequirementDescription
image_urlsRequired

String array with one or two public, directly downloadable URLs. The first image is Start; the optional second image is End. Preserve this order.

promptOptional

String. Directs action, camera movement, visual change, and sound intent after the start frame; 1–20,000 characters after trimming when provided.

durationOptional

Integer. Sets output length from 2 through 30 seconds, inclusive; default is 5.

resolutionOptional

String. Sets output resolution; default is 720p.

480p720p1080p
aspect_ratioOptional

String. Controls output framing; default is adaptive.

adaptive16:94:31:13:49:16
audioOptional

Boolean. Requests a generated audio track; default is true. The audio toggle does not change the credit rate.

truefalse
seedOptional

Integer. Optional reproducibility seed from 0 through 2147483647.

enable_safety_checkerOptional

Boolean. Enables the safety checker; default is true.

truefalse

How to Use

  1. Upload the start frameProvide a clear public image URL that already holds the subject, composition, lighting, and style you want at frame one.

  2. Add an end frame (optional)When the closing pose or product state needs a visual destination, add a second compatible image URL.

  3. Describe the motionWrite what happens after the still: subject action, camera path, lighting change, and sound intent.

  4. Set durationChoose an integer from 2 through 30 seconds; the default is 5 for first drafts.

  5. Choose resolutionSelect 480p for motion checks, or 720p / 1080p for review and delivery.

  6. Choose aspect ratioPick adaptive, 16:9, 4:3, 1:1, 3:4, or 9:16 to match the delivery framing.

  7. Configure audioKeep audio enabled for synchronized sound; turn it off for a silent clip.

  8. Generate the videoClick Run, then preview picture and sound together in the output area when the task finishes.

Pricing

Price depends only on output duration and resolution; the audio toggle does not change the rate.

UsageRateDetails
480p13.6 credits/output sec ($0.068/sec)2 seconds costs 27.2 credits ($0.136), 5 seconds costs 68 credits ($0.340), and 30 seconds costs 408 credits ($2.04).
720p28 credits/output sec ($0.14/sec)2 seconds costs 56 credits ($0.280), 5 seconds costs 140 credits ($0.70), and 30 seconds costs 840 credits ($4.20).
1080p56 credits/output sec ($0.28/sec)2 seconds costs 112 credits ($0.560), 5 seconds costs 280 credits ($1.40), and 30 seconds costs 1,680 credits ($8.40).

Best Use Cases

  • Product stills into motionTurn an approved product photo into a showcase clip with rotation, material movement, or a light change for launch teasers.

  • Poster and key visual animationAnimate a locked campaign still into a short motion asset for social, display, or presentation use.

  • Open-and-close transitionsConnect compatible first and last frames into a transition study for a reveal or composition change.

  • Character keyframe performanceUse a portrait or character still, then prompt expression, gesture, and camera response into a performance clip.

  • Concept art into shotsBring a concept painting into a short exploratory shot to evaluate motion and framing before production.

Pro Tips

  • Treat the start image as the opening sentence; use the prompt for what happens next instead of restating visible details.
  • Replace 'make the portrait move' with a visible progression: she looks toward the window, exhales, then turns back as the camera slowly pushes in.
  • When using a second image, keep identity, lighting logic, and art direction compatible across both stills.
  • Structure the prompt as subject action, scene and lighting, camera and shot, dialogue and sound, then timeline.
  • Validate motion at 480p / 5 seconds, then render 1080p and longer durations once the action plan holds.

Notes

  • image_urls requires 1–2 public http(s) URLs; the first is Start and the optional second is End.
  • Generation is asynchronous; retain task_id and stop tracking when the task reaches finished or failed.

Wan 3.0 Prime Image To Video API — Frequently asked questions

What is the Wan 3.0 Prime Image-to-Video API?

Wan 3.0 Prime Image-to-Video is an Alibaba Tongyi Lab model for generating high-definition video from images. It animates static starting frames into continuous takes up to 30 seconds at up to 1080p resolution with native audio, supporting an optional ending frame for precise end-state control. Built on Wan 3.0 Prime's upgraded spatiotemporal alignment architecture, it preserves source facial likeness, intricate apparel textures, and lighting environments while delivering enhanced physical motion dynamics. You can call it programmatically or try it from the playground above.

What motion improvements does Wan 3.0 Prime Image-to-Video provide over Wan 3.0?

Wan 3.0 Prime substantially improves character facial stability, joint articulation, and complex physical dynamics (such as hair sway, flowing water, and fabric motion) under significant camera movements, greatly reducing distortion when animating from static images.

How do I guide video transitions with an end frame in Wan 3.0 Prime Image-to-Video?

Provide two image URLs in the image_urls array, where the first acts as the start frame and the second defines the closing frame. The model calculates spatiotemporal interpolation between both compositions, generating smooth intermediate dynamics and camera shifts across your chosen duration.

Does Wan 3.0 Prime Image-to-Video synthesize synchronized sound for photos?

Yes. With joint audiovisual diffusion, the model automatically analyzes visual context and prompt descriptions to synthesize matching ambient acoustics, motion foley, and impact sounds without requiring external audio inputs.

How can I ensure subject identity consistency in Wan 3.0 Prime Image-to-Video?

Upload a clear, well-lit starting image with distinct facial or product features, and describe specific motion directions in the prompt while avoiding contradictory character changes. The model anchors identity features directly from the initial still.

Does Wan 3.0 Prime Image-to-Video support 9:16 vertical video?

Yes. When aspect_ratio is set to adaptive, the output matches the aspect ratio of the first input image. You can also explicitly specify 9:16 to adapt horizontal source stills into vertical video compositions optimized for mobile feeds.

Can Wan 3.0 Prime Image-to-Video generate motion without a text prompt?

Yes. If no prompt is provided, the model automatically infers plausible physical dynamics and camera drift based on the image's scene content. Adding prompt text allows you to direct explicit trajectories, character actions, and sound design.