Hailuo 2.3 Standard Text to Video API

minimax/hailuo-2.3/standard/text-to-video

Hailuo 2.3 Standard Text to Video transforms natural language prompts into fluid 768p video scenes, featuring 6-second or 10-second clips, realistic physical simulation, and expanded multi-style rendering across cinematic, anime, and CG aesthetics. It faithfully executes complex action directives and camera choreography while maintaining environmental coherence and subtle lighting depth.

Input
854/5000
OutputReady
768p · 10 seconds · 70 credits = $0.350
Continue with

Examples

One continuous ten-second side-view 2D cel animation, clean dark outlines and flat painted violet crystal-cave background. Exactly ONE small amber jelly creature with exactly two black eyes and no limbs moves from left to right. Its body is a single cohesive blob throughout. It meets a gap between two stationary crystal pillars that is only slightly narrower than its body. It gently squashes sideways to squeeze through the short gap, rounds out again on the other side, then makes two small hops to the right and settles. The same single face stays attached to the front of its body. Keep the entire creature visible. Moderate elastic deformation only: never stretch into a long string, never break apart, divide, duplicate or leave another blob behind. Slow sideways camera follows one character in one shot. No cuts, no text, no logo, no watermark.

A single uninterrupted documentary shot inside a transparent indoor skydiving wind tunnel. One adult beginner wearing a plain cobalt flight suit, clear goggles and a helmet floats horizontally above the safety mesh. Airflow gently rolls the body sideways; the flyer adjusts the angle of both arms and returns to a stable belly-down hover. Keep the entire body visible with believable joint motion, small corrective movements and fluttering fabric. Static medium-wide camera outside the glass, neutral indoor light, no other flyers. No cuts, no lettering, no logos, no advertising, no watermark.

A charming handcrafted stop-motion animated miniature scene in a dry eucalyptus woodland diorama. One original WOODEN ECHIDNA PUPPET investigates a small hole beneath a pale tree root. The puppet has a low rounded cork body, many short tapered wooden quills, a small dark carved snout, two bead eyes and short jointed wooden digging feet. It is visibly a stylized handmade animal character with wood grain and tiny hinge joints, NOT a live animal or scientific reconstruction. In one continuous six-second shot, it plants its rear feet, uses one forefoot to scrape a small pile of loose miniature red soil backward twice, then lowers its snout toward the hole and pauses. Small dirt crumbs move only where the foot touches. Keep exactly one puppet, stable crafted parts, clear foot-ground contact and simple deliberate stop-motion gestures. Fixed three-quarter camera, warm soft miniature-set lighting, dried eucalyptus leaves. No humans, labels, advertising, logos, captions or watermark.

Hailuo 2.3 Standard Text to Video

Hailuo 2.3 Standard Text to Video is developed by MiniMax for high-fidelity scene synthesis directly from natural language prompts. By expanding text prompt capacity up to 5,000 characters, creators can direct intricate multi-action sequences, detailed environmental lighting, and dynamic camera choreography. The model delivers continuous 6-second or 10-second video clips at 768p resolution, supporting realistic, anime, and illustration aesthetics with cost-effective per-video credit pricing.

Why Choose This?

  • Pure Text-Driven Scene CreationBuild vivid characters, environments, and motion directly from natural language without uploading starting assets.

  • Extended 5,000-Character Prompt CapacityDirect complex multi-stage narratives, sensory lighting, and specific camera paths with extensive text guidance.

  • Realistic Physical Motion SimulationAccurately replicates natural gravity, fluid dynamics, and bodily momentum for believable physical interactions.

  • Multi-Style Visual VersatilitySupports photorealistic scenes alongside anime, digital illustration, and game CG styles with consistent aesthetic rendering.

  • Predictable Per-Video BillingClear pricing of 35 credits for 6s or 70 credits for 10s, with automatic credit refunds if a generation task fails.

Parameters

ParameterRequirementDescription
promptRequired

Required nonblank string, trimmed before validation. Maximum 5,000 Unicode characters.

durationOptional

6 or 10 seconds; defaults to 6.

Default6
resolutionOptional

Fixed to 768p for this endpoint; used when omitted.

Default768p
prompt_optimizerOptional

Optional boolean; omit to leave unspecified upstream. No API default. The playground starts with false; no extra charge.

How to Use

  1. Define Subject and SettingEstablish character identity, focal props, and atmospheric environment in the prompt to ground the composition.

  2. Choreograph Actions and Camera MotionDetail sequential character actions and explicit camera mechanics like panning, tracking, or dolly zoom.

  3. Select DurationChoose 6 seconds for dynamic concise clips or 10 seconds for extended narrative progression at fixed 768p resolution.

  4. Configure Prompt OptimizerEnable prompt_optimizer when working from concise prompts to enrich cinematic lighting and environmental details at no extra cost.

  5. Submit and Retrieve VideoDispatch the asynchronous task and poll status using task_id to download the completed MP4 video file.

Pricing

1 credit = $0.005. Billed per video; prompt optimization does not change the rate.

UsageRateDetails
768p / 6s35 credits ($0.175)Per video
768p / 10s70 credits ($0.350)Per video

Best Use Cases

  • Cinematic Concept PrototypingTransform script scenes into dynamic video mockups to evaluate camera movement and pacing before physical production.

  • Anime and Stylized Content CreationGenerate stylized animations and illustrative sequences with frame-to-frame stylistic consistency.

  • Social Media and Digital CampaignsProduce eye-catching realistic video clips and visual loops for social feeds and promotional channels.

  • Commercial Motion StoryboardingVisualize product showcases, dynamic environments, and storytelling vignettes directly from descriptive ad copy.

Pro Tips

  • Decouple Subject Motion from Camera Direction: Describe what characters do and how the camera moves in distinct sentences for clearer execution.
  • Take Advantage of 5,000 Characters: Include sensory descriptions covering lighting angle, weather ambiance, surface reflections, and pacing.
  • Strategize Prompt Optimizer Usage: Turn prompt_optimizer on for brief conceptual prompts, and keep it off when strict adherence to precise directorial phrasing is necessary.
  • Select Duration Based on Scene Complexity: Use 6-second clips (35 credits) for single distinct actions, and 10-second clips (70 credits) for multi-phase narrative changes.
  • Anchor Abstract Style with Concrete Physics: Instead of generic adjectives like 'dramatic', describe physical cues such as 'shadows stretch across damp concrete' or 'cloth billows in crosswinds'.

Notes

  • Text-Only Input Interface: This endpoint generates video exclusively from the prompt parameter; image attachments are not accepted.
  • Fixed 768p Output Specification: Video resolution is set to 768p, with integer duration options of 6 or 10 seconds.
  • Asynchronous Execution and Credit Guarantee: Tasks run asynchronously via unique task_id; credits are deducted upon submission and refunded automatically if execution fails.

Hailuo 2.3 Standard Text to Video API frequently asked questions

What is the Hailuo 2.3 Standard Text to Video API?

Hailuo 2.3 Standard Text to Video is a MiniMax model for video generation from text. It generates continuous dynamic videos at 768p resolution directly from text prompts, supporting 6-second or 10-second durations, realistic physical simulation, and optional prompt refinement. Built on MiniMax's advanced video generation architecture, it faithfully executes narrative choreography and camera movement while preserving scenic coherence and lighting fidelity across realistic, anime, and CG art styles. You can call it programmatically or try it from the playground above.

Does Hailuo 2.3 Standard Text to Video support 10-second generations?

Yes. The model provides discrete 6-second and 10-second duration options, with 6 seconds as the default. Choosing 10 seconds allows creators to depict multi-stage actions, extended narrative arcs, and gradual camera transitions in a single clip.

How does the prompt optimizer work in Hailuo 2.3 Standard Text to Video?

The prompt optimizer (prompt_optimizer) enriches concise user prompts by automatically incorporating cinematic lighting, realistic textures, and camera framing details. It is an optional boolean parameter disabled by default in the playground, and enabling it incurs no extra credits.

What prompt length is supported by Hailuo 2.3 Standard Text to Video?

The prompt parameter supports up to 5,000 Unicode characters after trimming leading and trailing whitespace. This extensive capacity enables creators to provide comprehensive shot lists, lighting cues, and character timing instructions in a single prompt.

How is Hailuo 2.3 Standard Text to Video billed?

Billing is calculated on a fixed per-video basis (1 credit = $0.005). A 6-second 768p video costs 35 credits ($0.175), while a 10-second 768p video costs 70 credits ($0.350). Enabling prompt optimization does not change the credit rate.

Can Hailuo 2.3 Standard Text to Video render diverse artistic styles?

Yes. In addition to photorealistic live-action scenes, the model features strong aesthetic representation for anime, digital illustration, ink wash painting, and game CG styles, maintaining visual coherence throughout the animation.

When should creators choose Hailuo 2.3 Standard over Pro?

The Standard tier is ideal for projects requiring 768p resolution, flexible duration choices between 6 and 10 seconds, and economical iteration starting at 35 credits per video. If your production requires native 1080p Full HD resolution for high-end cinematic deliverables, consider Hailuo 2.3 Pro Text to Video.