Hailuo 02 Standard Text to Video API

minimax/hailuo-02/standard/text-to-video

Hailuo 02 Standard Text to Video transforms text prompts into cinematic 512P and 768P video scenes, supporting 6-second or 10-second clips, realistic physical simulation, and optional prompt refinement. It preserves compositional spatial coherence and fine lighting texture while delivering fluid camera movement and expressive character dynamics.

Input
530/1500
OutputReady
768P · 6 seconds · 42 credits = $0.210

Examples

A single continuous cinematic side-tracking shot beneath a wide concrete urban overpass in soft afternoon light. One adult skateboarder in a rust-orange jacket and dark trousers rides smoothly up a very low concrete bank, briefly clears the lip with the skateboard beneath both feet, lands on all four wheels with bent knees, and rolls forward. Keep the whole body and board visible. Foreground bridge columns pass slowly across the edge of the frame with strong parallax. Realistic human anatomy, balanced landing and grounded wheel contact. No cuts, no lettering, no logos, no advertising.

One uninterrupted ten-second close tracking shot in a quiet old underground print workshop with cool window light. An antique mechanical printing press starts slowly. Interlocking brass gears rotate together, driving a heavy ink roller and steadily feeding one sheet of paper through the press. The sheet emerges onto the receiving tray carrying a simple abstract blue-and-red woodcut pattern with absolutely no letters. The camera moves from the visible gears to the outgoing sheet without a cut. Physically connected mechanisms, consistent paper thickness, no people, no branding, no writing.

Hailuo 02 Standard Text to Video

Hailuo 02 Standard Text to Video is developed by MiniMax for high-fidelity scene synthesis from natural language prompts. By describing subject appearance, setting dynamics, and camera choreography, creators can generate continuous 6-second or 10-second video clips. The model simulates natural physical interactions and nuanced motion across 512P and 768P resolutions with transparent per-second pricing.

Why Choose This?

  • Pure Text-Driven Scene CreationBuild vivid characters, environments, and motion directly from natural language without uploading starting assets.

  • Realistic Physical Motion SimulationAccurately replicates gravity, fluid inertia, and soft-body collisions for believable environmental interactions.

  • Flexible Duration and Resolution TiersSupports 6-second agile scenes or 10-second extended shots across economical 512P and default 768P resolutions.

  • Optional Intelligent Prompt OptimizationExpands concise prompts with rich cinematography, textural lighting, and environmental nuances without extra charges.

  • Predictable Per-Second BillingCharges strictly by output resolution and duration, with automatic credit refunds if a task encounters an error.

Parameters

ParameterRequirementDescription
promptRequired

Required nonblank string, trimmed before validation. Maximum 1500 Unicode characters.

resolutionOptional

512P or 768P; defaults to 768P.

Default768P
durationOptional

6 or 10 seconds; defaults to 6.

Default6
prompt_optimizerOptional

Optional boolean; omit to leave unspecified upstream. The playground starts with false.

How to Use

  1. Define Subject and SettingEstablish character appearance, focal props, and scenic backdrops in the opening sentence to anchor the composition.

  2. Outline Motion and Camera MotionDescribe actions chronologically and specify explicit camera directions such as dolly zoom, panning, or tracking.

  3. Select Resolution and DurationPick 6 seconds for concise clips or 10 seconds for narrative progression, alongside 512P or 768P resolution.

  4. Toggle Prompt OptimizationEnable prompt_optimizer when working from concise concepts to enrich cinematic atmosphere and lighting.

  5. Submit and Retrieve ResultDispatch the asynchronous generation task and poll with task_id to stream or download the finished MP4 video file.

Pricing

1 credit = $0.005. Billed by resolution and output duration.

UsageRateDetails
512P / 6s18 credits ($0.090)3 credits/second
512P / 10s30 credits ($0.150)3 credits/second
768P / 6s42 credits ($0.210)7 credits/second
768P / 10s70 credits ($0.350)7 credits/second

Best Use Cases

  • Commercial Storyboard PrototypingTranslate script lines into dynamic motion previews to validate pacing and camera angles before physical production.

  • Social Media & Viral ContentQuickly generate captivating realistic video clips and aesthetic visual loops for digital channels.

  • Film and Drama PrevisualizationTurn narrative script beats into continuous 6-to-10-second dramatic scenes to guide directorial vision.

  • World-Building Concept VisualizationBreathe life into speculative world descriptions, sci-fi machinery, and natural phenomena.

Pro Tips

  • Decouple Subject Motion from Camera Direction: Specify what the subject does and how the camera moves in distinct sentences for cleaner choreography.
  • Leverage the 1500-Character Prompt Capacity: Use detailed sensory adjectives covering atmospheric haze, reflective textures, and depth of field.
  • Choose When to Enable Prompt Optimization: Turn the optimizer on for concise conceptual ideas, and keep it off when strict adherence to a precise directorial prompt is required.
  • Test Pacing with 512P First: Validate action timing quickly with 512P / 6s (18 credits) before committing to 768P / 10s for final deliverables.
  • Anchor Abstract Adjectives to Physical Cues: Replace words like 'epic' or 'amazing' with concrete physics cues such as 'dust kicks up underfoot' or 'light reflects off wet asphalt'.

Notes

  • Text-Only Input Contract: This endpoint accepts a single trimmed prompt string up to 1500 Unicode characters without image attachments.
  • Discrete Resolution and Duration Constraints: Parameters only accept 512P or 768P for resolution, and 6 or 10 integer seconds for duration.
  • Asynchronous Execution & Safe Billing: Tasks process asynchronously via unique task_id; credits are deducted upon validation and refunded immediately if execution fails.

Related Models

Hailuo 02 Standard Text to Video API frequently asked questions

What is the Hailuo 02 Standard Text to Video API?

Hailuo 02 Standard Text to Video is a MiniMax model for video generation from text. It generates continuous dynamic videos at 512P or 768P resolution directly from text prompts, supporting 6-second or 10-second durations, realistic physical simulation, and optional prompt refinement. Built on MiniMax's advanced video generation architecture, it faithfully executes narrative choreography and camera movement while preserving scenic coherence and lighting fidelity. You can call it programmatically or try it from the playground above.

Does Hailuo 02 Standard Text to Video support 10-second generations?

Yes. The model provides discrete 6-second and 10-second duration options, with 6 seconds as the default. Choosing 10 seconds allows creators to depict multi-stage actions, extended narrative arcs, and gradual camera transitions in a single clip.

How does the prompt optimizer work in Hailuo 02 Standard Text to Video?

The prompt optimizer (prompt_optimizer) enriches concise user prompts by automatically incorporating cinematic lighting, realistic textures, and camera framing details. It is an optional boolean parameter disabled by default in the playground, and enabling it incurs no extra credits.

What prompt length is supported by Hailuo 02 Standard Text to Video?

The prompt is trimmed of leading and trailing whitespace and accepts between 1 and 1,500 Unicode characters. This generous length allows detailed descriptions of subject characteristics, progressive actions, and environmental lighting.

How is Hailuo 02 Standard Text to Video billed?

This endpoint is billed on a per-second basis based on output resolution and duration (1 credit = $0.005). The 512P tier costs 3 credits/second (18 credits for 6s, 30 credits for 10s), while 768P costs 7 credits/second (42 credits for 6s, 70 credits for 10s).

Can Hailuo 02 Standard Text to Video control camera movement via prompts?

Yes. You can incorporate standard cinematography terms (such as 'slow push-in', 'lateral tracking shot', or 'wide high-angle crane') directly into your prompt. The model coordinates camera movement seamlessly with the subject's actions.

When should I choose Hailuo 02 Standard Text to Video over the Pro version?

Choose the Standard mode when you require explicit 6-second or 10-second duration control, economical per-second pricing (such as 512P at 3 credits/sec), or high-throughput conceptual prototyping. Choose the Pro mode when producing final hero assets demanding heightened cinematic dynamics.