Hailuo 2.3 Standard Image to Video API

minimax/hailuo-2.3/standard/image-to-video

Hailuo 2.3 Standard Image to Video animates static starting images into expressive 768p video scenes, supporting 6-second or 10-second outputs, natural facial micro-expressions, and responsive motion prompt control. It preserves original character identity, compositional framing, and lighting atmosphere while introducing fluid physical movement and cinematic camera transitions.

Input
545/5000
Jumping Spider Notices You

Required first frame. Playground uploads support JPG, PNG and WebP, up to 10 MiB. Use JSON mode to supply an HTTP(S) URL.

OutputReady
768p · 6 seconds · 35 credits = $0.175
Continue with

Examples

Animate the exact jumping spider in this first frame. In one uninterrupted macro shot it takes two short sideways steps on the observation surface, rotates its body to face the camera, lifts its front pair of legs and holds still. Preserve its eight legs, large forward-facing eyes, compact body and detailed hairs without adding appendages. Natural tiny foot contacts and deliberate pauses. Keep the camera, background and lighting fixed with enough depth of field to see the body. No cuts, no lettering, no logos, no advertising, no watermark.

Animate this exact first frame in one continuous realistic shot. The single adult mover holds the same flexible mattress vertically at the narrow apartment doorway. He pauses, rotates the mattress slightly onto its diagonal edge, turns his shoulders sideways and takes a short step through the opening with it. The mattress bends only slightly where it brushes the door frame; both hands keep a consistent grip. Show believable weight transfer and continuous contact. Keep the mover, clothing, mattress and doorway unchanged. Fixed medium-wide camera; no cuts or extra people. No cuts, no lettering, no logos, no advertising, no watermark.

Animate this original black-and-white rubber-hose cartoon in one uninterrupted ten-second shot. The round inkwell character first scrunches its face and leans backward trying to suppress a sneeze, then sneezes with a short elastic squash and stretch. Its one loose lid pops straight upward, falls back onto its own opening and settles with one small bounce. The bottle body wobbles and gradually becomes still. Keep the same character, black ink fill, small arms, legs and checkerboard floor throughout. The lid is never duplicated. Fixed camera, vintage hand-inked animation, no color. No cuts, no lettering, no logos, no advertising, no watermark.

Hailuo 2.3 Standard Image to Video

Hailuo 2.3 Standard Image to Video is developed by MiniMax to transform still images into fluid, lifelike video sequences. Creators provide a single high-quality starting image URL alongside natural language prompts to animate characters, environments, and objects. The model generates continuous 6-second or 10-second video clips at 768p resolution, reproducing nuanced facial micro-expressions and physically accurate motions while preserving input composition with transparent credit pricing.

Why Choose This?

  • High-Fidelity Subject Likeness PreservationMaintains facial contours, costume textures, and key visual attributes faithfully across the entire animation sequence.

  • Subtle Facial Micro-Expression ModelingAccurately conveys nuanced emotional shifts such as subtle smiles, thoughtful glances, and expressive character moments.

  • Flexible 6-Second and 10-Second DurationsSupports 6-second focused action clips or 10-second extended shots for comprehensive narrative development.

  • Responsive Motion Directive ExecutionFaithfully translates detailed textual instructions for body actions, environmental dynamics, and camera movement.

  • Economical and Transparent Credit PricingFixed pricing at 35 credits for 6s or 70 credits for 10s, with automatic credit refunds if task execution fails.

Parameters

ParameterRequirementDescription
promptRequired

Required nonblank string, trimmed before validation. Maximum 5,000 Unicode characters.

durationOptional

6 or 10 seconds; defaults to 6.

Default6
resolutionOptional

Fixed to 768p for this endpoint; used when omitted.

Default768p
start_image_urlRequired

Required first-frame HTTP(S) image URL with a hostname and no embedded credentials. End frames, arrays and data URLs are not supported.

prompt_optimizerOptional

Optional boolean; omit to leave unspecified upstream. No API default. The playground starts with false; no extra charge.

How to Use

  1. Host and Provide Starting ImageUpload or host a high-resolution JPG, PNG, or WebP image and provide its accessible HTTP(S) URL in start_image_url.

  2. Direct Movement and Camera MotionWrite prompts specifying subject physical actions, environmental reactions, and camera moves such as pan or zoom.

  3. Select Clip DurationPick 6 seconds for swift dynamic actions or 10 seconds for detailed narrative progression at fixed 768p resolution.

  4. Configure Prompt OptimizerOptionally activate prompt_optimizer to enrich motion cues and cinematic lighting nuances without extra credits.

  5. Submit Task and Retrieve VideoDispatch the asynchronous task and poll status using task_id to obtain the final MP4 video download URL.

Pricing

1 credit = $0.005. Billed per video; prompt optimization does not change the rate.

UsageRateDetails
768p / 6s35 credits ($0.175)Per video
768p / 10s70 credits ($0.350)Per video

Best Use Cases

  • Portrait and Character AnimationBreathe dynamic life into digital portraits, gaming characters, and avatars while maintaining character likeness.

  • E-Commerce Product DemonstrationsAnimate static product photography into engaging commercial videos showcasing materials, reflections, and motion.

  • Anime and Illustration DynamicsTurn 2D anime illustrations and hand-drawn concept art into fluid animated scenes with stylistic consistency.

  • Storyboard Frame-to-Scene RealizationTransform keyframe illustrations into full animated sequences to validate pacing, action timing, and blocking.

Pro Tips

  • Anchor Prompts to the Starting Pose: Describe actions originating naturally from the subject's pose in the starting image for smoother transitions.
  • Detail Micro-Expressions and Gaze: Explicitly describe subtle shifts in facial expressions and eye contact to produce compelling close-up performances.
  • Use Active Direction Verbs: Employ precise action verbs like 'turns head toward camera', 'walks through falling leaves', or 'camera slowly tracks right'.
  • Choose Duration to Match Action Scale: Use 6-second clips (35 credits) for single physical motions, and 10-second clips (70 credits) for multi-beat sequences.
  • Ensure High-Quality Source Lighting: Clear, well-lit starting images allow the model to accurately deduce realistic shadows and reflective highlights during motion.

Notes

  • Single Starting Frame Interface: This endpoint requires a single accessible HTTP(S) image URL in start_image_url; end frames and image arrays are not accepted.
  • Fixed 768p Resolution Contract: Video generation outputs fixed 768p resolution, with integer duration options of 6 or 10 seconds.
  • Asynchronous Execution and Credit Protection: Tasks process asynchronously via unique task_id; credits are deducted upon submission and refunded immediately if execution fails.

Hailuo 2.3 Standard Image to Video API frequently asked questions

What is the Hailuo 2.3 Standard Image to Video API?

Hailuo 2.3 Standard Image to Video is a MiniMax model for video generation from images. It animates static images into continuous 768p video scenes based on starting image inputs and text prompts, supporting 6-second or 10-second clips, subtle facial micro-expressions, and physical motion simulation. Built on MiniMax's enhanced multimodal video architecture, it preserves character identity, compositional framing, and lighting texture while generating natural body movements and cinematic camera transitions. You can call it programmatically or try it from the playground above.

What starting image formats are supported by Hailuo 2.3 Standard Image to Video?

The endpoint accepts publicly accessible HTTP(S) URLs pointing to JPG, PNG, or WebP images via the start_image_url parameter. In the interactive playground, images up to 10 MiB can be uploaded directly.

Does Hailuo 2.3 Standard Image to Video support an end frame?

This endpoint focuses strictly on start-frame guided generation through the start_image_url parameter. To steer the conclusion of your video, describe the closing action, subject placement, and final camera framing directly within the prompt.

How does Hailuo 2.3 Standard Image to Video preserve character facial likeness?

The model extracts identity and textural features from the starting frame to maintain facial anatomy, hair details, and costume consistency across the entire clip, allowing natural head turns and micro-expressions without facial warping.

Can Hailuo 2.3 Standard Image to Video generate 10-second animations?

Yes. The endpoint supports discrete duration options of 6 seconds (default, 35 credits) and 10 seconds (70 credits). Selecting 10 seconds provides ample temporal span for extended character dialogue gestures and multi-angle scene choreography.

How is Hailuo 2.3 Standard Image to Video billed?

Billing is calculated on a fixed per-video basis (1 credit = $0.005). A 6-second 768p video costs 35 credits ($0.175), while a 10-second 768p video costs 70 credits ($0.350). Enabling prompt optimization incurs no extra charges.

When should creators choose Hailuo 2.3 Standard Image to Video over Pro?

The Standard tier is recommended when you need flexible 6-second or 10-second durations, reliable 768p output, and economical per-video rates for rapid asset iteration. Choose Hailuo 2.3 Pro Image to Video when your production requires native 1080p Full HD fidelity for commercial showcase pieces.