Wan 3.0 Image-to-Video API

alibaba/wan-3.0/image-to-video

Wan 3.0 (Image-to-Video) transforms start-frame images and optional prompts into dynamic video, with optional end-frame guidance and 2 to 30 second continuous generation. It preserves source subject identity, texture, and visual composition while introducing physically consistent motion and synchronized native audio.

Input

583/20000
Start frame preview
End frame preview

Output

Ready
5 sec × $0.100/sec = $0.500

Continue with

Examples

Begin exactly from Image 1 and finish on Image 2. In one continuous five-second shot, the suited tabby cat leans forward with mock seriousness, then suddenly face-plants into the keyboard. Papers twitch. The coffee cup stays put. Camera: locked medium shot with a tiny downward tilt at the impact. No cuts. Preserve the cat's face markings, suit, tie, desk, and lighting. Synchronized audio: chair creak, rapid keyboard clacks, a muffled meow. No readable text, letters, numbers, captions, labels, logos, brands, watermarks, advertisements, posters, UI screens, or product packaging.

Start from Image 1. In one continuous five-second shot, the stone frog statue slowly blinks both eyes once, then lifts the tiny sunglasses from the pedestal and slides them onto its face with a self-satisfied pause. Camera: locked medium shot, very slow push-in. No cuts. Keep the courtyard, moss, and stone texture consistent. Synchronized audio: faint stone grit, a soft blink, sunglasses clicking onto stone. No readable text, letters, numbers, captions, labels, logos, brands, watermarks, advertisements, posters, UI screens, or product packaging.

Start from Image 1. In one continuous five-second shot, the refrigerator door stays slightly open while the leftover containers sway gently left and right in a shy little queue, as if waiting their turn. Camera: locked interior shot, almost still. No cuts. Synchronized audio: low fridge hum, soft plastic taps. No readable text, letters, numbers, captions, labels, logos, brands, watermarks, advertisements, posters, UI screens, or product packaging.

Wan 3.0 Image-to-Video

Wan 3.0 Image-to-Video generates continuous high-definition video with native synchronized audio from an initial image and optional text prompt. It faithfully preserves subject appearance, surface texture, and framing while synthesizing natural physical dynamics, camera movements, and environmental soundscapes.

Why Choose This?

  • Robust subject fidelityAnchor facial features, hair details, and product textures steadily over takes up to 30 seconds without visual distortion or character drift.

  • Optional end-frame trajectory guidanceSupply both opening and closing images to guide smooth, physically plausible transitions between specific poses or compositions.

  • Multimodal motion and camera controlPair input pictures with descriptive text prompts to orchestrate precise pans, tilts, lighting shifts, and subtle facial micro-expressions.

  • Native synchronized soundscapesSynthesize authentic action foley and ambient soundscapes alongside visual animation directly through the underlying joint diffusion model.

  • Flexible resolution tiersRender outputs at 480p, 720p, or 1080p, automatically preserving the source image aspect ratio or matching standard commercial formats.

Parameters

ParameterRequirementDescription
image_urlsRequired

Array of strings. 1 to 2 publicly accessible image HTTP(S) URLs; first image acts as the start frame, optional second image acts as the end frame. Maximum 30MB per image, supporting JPEG, PNG, and WebP.

promptOptional

String. Guides camera movement, subject action, and audio design; supports 1 to 20,000 characters.

durationOptional

Integer. Output duration in whole seconds between 2 and 30; playground defaults to 5 seconds.

Default5
resolutionOptional

String. Native output resolution tier; choices include 480p, 720p (default), or 1080p.

Default720p480p1080p
aspect_ratioOptional

String. Framing ratio; defaults to adaptive (inherits input image ratio), or accepts 16:9, 4:3, 1:1, 3:4, and 9:16.

Defaultadaptive16:94:31:13:49:16
audioOptional

Boolean. Determines whether to synthesize a synchronized audio track alongside the video; defaults to true at no extra cost.

Defaulttruefalse
seedOptional

Integer. Random seed between 0 and 2,147,483,647 for reproducible trajectories and dynamics.

enable_safety_checkerOptional

Boolean. Enables automated safety filtering on prompts and generated outputs; defaults to true.

Defaulttruefalse

How to Use

  1. Upload a clear start-frame imageSelect a well-lit image with distinct subjects as the opening frame (up to 30MB, recommended 720p or higher resolution).

  2. Optionally attach an end-frame imageIf you require the clip to resolve to a predetermined composition or pose, provide a second image to guide the final frame.

  3. Describe action and camera movementAdd concise instructions detailing character movement and camera trajectory (e.g., The woman turns slowly towards the lens with a gentle smile as the camera pushes in).

  4. Configure duration and resolutionAdjust the duration slider between 2 and 30 seconds, select 720p or 1080p, and leave ratio as adaptive to match your source image.

  5. Check audio and advanced optionsKeep audio enabled to generate matched sound effects and room ambience, or lock seed for reproducible motion studies.

  6. Submit and inspect playbackRun the task to start asynchronous processing, then review real-time render progress and preview or download the completed MP4 video.

Pricing

Wan 3.0 Image-to-Video charges by generated output second based strictly on the selected resolution tier; enabling or disabling audio carries no extra fee (1 credit = $0.005).

UsageRateDetails
480p10 credits / output sec ($0.05 / sec)Standard definition tier. 5-second default is 50 credits ($0.25); 30-second maximum is 300 credits ($1.50).
720p (Default)20 credits / output sec ($0.10 / sec)High definition tier. 5-second default is 100 credits ($0.50); 30-second maximum is 600 credits ($3.00).
1080p40 credits / output sec ($0.20 / sec)Full high definition flagship tier. 5-second default is 200 credits ($1.00); 30-second maximum is 1,200 credits ($6.00).

Best Use Cases

  • E-commerce product showcasesConvert static product photography into cinematic showcase clips with realistic lighting shifts and subtle camera glides.

  • Portrait and character animationBreathe life into illustrations, game concept art, or portrait photos with natural eye blinks, hair sway, and facial expressions.

  • Historical archival reanimationTurn vintage photos and landscape captures into dynamic historical vignettes accompanied by authentic atmospheric room tone.

  • Storyboard keyframe interpolationConnect opening and closing storyboard frames with smooth, physics-informed motion to preview scene transitions.

Pro Tips

  • Source high-clarity opening frames: The quality of the initial image directly determines output fidelity; prefer sharp images with defined lighting and high contrast.
  • Match character style between keyframes: When using an end frame, keep subject identity, wardrobe, and illumination consistent across both images for natural interpolation.
  • Focus prompts on incremental motion: Since subject appearance is already defined by the image, focus your text on verbs describing motion dynamics and camera paths.
  • Retain adaptive aspect ratio: Unless targeting a specific social format, adaptive prevents unnecessary framing crops or stretching on non-standard source images.
  • Scale duration with motion complexity: Subtle expressions work well within 4 to 6 seconds, whereas sweeping physical actions benefit from 10 to 15 seconds.

Notes

  • Input image limits: Provide exactly 1 start-frame image, with an optional 2nd end-frame image; maximum 2 images per request.
  • File format requirements: Images must be publicly reachable URLs under 30MB each in standard JPEG, PNG, or WebP format.
  • Whole-second duration input: The duration parameter accepts whole integers between 2 and 30 seconds.

Wan 3.0 Image-to-Video API — Frequently Asked Questions

What is the Wan 3.0 Image-to-Video API?

Wan 3.0 Image-to-Video is an Alibaba Tongyi Lab model for generating video from images. It animates static starting images—with optional ending frame guidance—into continuous videos up to 30 seconds at up to 1080p resolution with native synchronized audio. Built on Diffusion Transformer and Wan-VAE 3D spatiotemporal architectures, it faithfully preserves the source subject's facial identity, clothing textures, and lighting while introducing smooth physical dynamics. You can call it programmatically or try it from the playground above.

How do I specify an end frame in Wan 3.0 Image-to-Video?

Provide a second public image URL in the image_urls array as your closing keyframe. The model computes geometric and lighting displacements between both stills to synthesize organic physical transitions and camera maneuvers across your selected duration.

Can Wan 3.0 Image-to-Video preserve clothing textures and facial likeness accurately?

Yes. The model anchors subject features, outfit fabric folds, and ambient lighting directly from the first frame. Adding explicit trajectory instructions or camera angles in the prompt ensures that subjects remain visually consistent as action unfolds.

Will Wan 3.0 Image-to-Video generate audio when using only a single static image?

Yes. With joint audiovisual diffusion, keeping audio enabled prompts the model to interpret visual cues and prompt actions, automatically synthesizing matching environmental foley and ambient acoustics without requiring manual sound uploads.

Does Wan 3.0 Image-to-Video support 9:16 vertical video generation?

Yes. When aspect_ratio is set to adaptive, the output matches the aspect ratio of the first input image. If your starting still is horizontal, you can also explicitly choose 9:16 to adapt the framing for mobile vertical distribution.

Does motion distort or degrade during a 30-second take in Wan 3.0 Image-to-Video?

No. The model leverages advanced spatiotemporal modeling to sustain physical realism across takes up to 30 seconds. For ambitious camera paths, we recommend describing gradual pacing and staged action beats in your prompt rather than sudden extreme shifts.

When having a starting frame, should I use Wan 3.0 Image-to-Video or Reference-to-Video?

Choose Image-to-Video if your video's opening shot must lock pixel-for-pixel onto the composition and camera framing of your image. Choose Reference-to-Video if you only need to borrow a character likeness or prop while creating a completely new opening camera setup and environment.