Hailuo 2.3 Pro Text to Video API

minimax/hailuo-2.3/pro/text-to-video

Hailuo 2.3 Pro Text to Video turns descriptive text prompts into cinematic native 1080p video clips with true-to-life physics simulation, high dynamic range lighting, and refined motion details. It accurately follows intricate multi-subject choreography, fluid camera pans, and natural inertia while delivering studio-grade visual clarity across fixed 6-second generations.

Input
946/5000
OutputReady
1080p · 6 seconds · 60 credits = $0.300
Continue with

Examples

A restrained cinematic comedy shot of ONE muscular adult amateur boxer reclining in a dental examination chair, wearing a plain dark charcoal T-shirt. Tight close-up of his face, neck and upper chest; his arms and all hands remain OUTSIDE the frame throughout. A single small dental inspection mirror on a thin metal stem enters slowly from the far left edge, held by an unseen dentist. It stays near the image edge and never touches the man. He notices it with only his eyes, swallows once with a visible small throat movement, then tightens his lips and furrows his brow while trying to seem brave. His head remains nearly still. Natural skin texture and subtle coherent facial muscle movement, consistent face, softly blurred clinical room background, fixed camera and soft light. One continuous six-second take, no cuts. No gloves, hands or extra people in frame, no mouth interior, no treatment, no blood, no writing, no logos, no watermark.

A single uninterrupted close-up of a quiet physics experiment on a dark laboratory bench. One large iridescent soap bubble is held between two clean vertical transparent acrylic plates. The right plate moves very slowly toward the left, gently flattening the bubble into a wider oval without bursting it; it then retreats slightly and the elastic soap film rounds out again. A thin wet contact boundary remains visible on each plate. Fine rainbow interference bands flow across the film, realistic transparent reflections and delicate oscillations. Locked camera, black background, soft rectangular laboratory lights; no people or labels. No cuts, no lettering, no logos, no advertising, no watermark.

Hailuo 2.3 Pro Text to Video

Hailuo 2.3 Pro Text to Video is MiniMax's flagship text-to-video generation model, engineered specifically for high-end cinematic visualization and commercial production. Operating at native 1080p full HD resolution with a fixed 6-second runtime, it delivers studio-quality spatial fidelity, physically accurate fluid and gravitational dynamics, and responsive camera motion control. Creators can input rich prompts up to 5,000 characters, utilize the complimentary prompt optimizer for lighting expansion, and integrate high-throughput rendering via a predictable 60-credit per video pricing tier.

Why Choose This?

  • Native 1080p Full HD FidelityRenders crisp 1080p resolution directly without post-upscaling artifacts, preserving intricate skin textures, garment folds, and environmental details.

  • Advanced Real-World Physics SimulationAccurately computes complex inertial motion, fluid splash dynamics, fabric drapery, and gravitational interactions across every frame.

  • Cinematic Lighting and Surface ReflectionsSimulates realistic volumetric scattering, specular highlights, shadow soft-falloff, and atmospheric perspective for true filmic aesthetics.

  • Extensive 5,000-Character Prompt CapacitySupports comprehensive multi-layer scene descriptions, directing lighting setups, shot progression, pacing, and subject interactions with precision.

  • Predictable Studio-Grade PricingFixed rate of 60 credits ($0.300) per 6-second full HD generation ensures straightforward budgeting for commercial production pipelines.

Parameters

ParameterRequirementDescription
promptRequired

Required nonblank string, trimmed before validation. Maximum 5,000 Unicode characters.

durationOptional

Only 6 seconds; defaults to 6.

Default6
resolutionOptional

Fixed to 1080p for this endpoint; used when omitted.

Default1080p
prompt_optimizerOptional

Optional boolean; omit to leave unspecified upstream. No API default. The playground starts with false; no extra charge.

How to Use

  1. Craft Detailed Scene PromptDraft a comprehensive description specifying subject appearance, physical action sequences, lighting style, and camera trajectory up to 5,000 characters.

  2. Confirm Resolution and Duration SpecificationsOutput is natively configured at 1080p resolution and a focused 6-second runtime for maximum per-frame cinematic rendering quality.

  3. Toggle Prompt OptimizerOptionally enable prompt_optimizer to allow upstream AI to expand shorthand prompts with professional lighting and cinematography cues at no added cost.

  4. Submit Async Generation RequestSend your POST request to /api/generate/submit with your API key to initialize rendering and receive a unique task_id immediately.

  5. Poll Status and Download VideoCheck the status endpoint until marked finished, then retrieve the hosted MP4 file URL from data.files for post-production or streaming.

Pricing

1 credit = $0.005. Billed per video; prompt optimization does not change the rate.

UsageRateDetails
1080p / 6s60 credits ($0.300)Per video

Best Use Cases

  • Commercial Advertising and Brand FilmsCreate broadcast-ready product showcases, visual effects, and high-impact hero shots with native 1080p clarity and premium lighting.

  • Film and Series Pre-VisualizationTranslate script excerpts and director notes into realistic dynamic camera previs shots, validating shot composition before live filming.

  • Game Cinematics and Teaser TrailersRender dynamic cutscenes, character introductions, and atmospheric environmental sequences with intense action and physics.

  • Social Media Brand CampaignsProduce captivating, high-fidelity promotional clips tailored for luxury, fashion, automotive, and technology brand narratives.

Pro Tips

  • Direct camera motion explicitly: Incorporate professional camera commands such as 'slow push-in', 'dolly left to reveal', or 'handheld tracking shot' to enhance visual rhythm.
  • Detail multi-layered environmental lighting: Describe distinct light sources like 'golden hour rim lighting with soft ambient blue fill' to maximize 1080p surface detail.
  • Structure actions chronologically: Describe sequential beats within the 6-second duration (e.g., 'the athlete pauses, looks up, then sprints forward') for cohesive timing.
  • Leverage prompt optimizer for shorthand concepts: Enable prompt_optimizer when testing rapid creative concepts to automatically fill in filmic stylistic nuances.
  • Specify material and atmospheric textures: Include physical descriptors such as 'mist swirling in damp cobblestone reflections' to activate the advanced physics engine.

Notes

  • Fixed 6-second duration design: Hailuo 2.3 Pro Text to Video is purposefully optimized for 6-second high-density sequences; requesting duration other than 6 is rejected.
  • Native 1080p rendering standard: Outputs render natively at full HD 1080p resolution without spatial interpolation, ensuring pristine professional clarity.
  • Async task lifecycle and credit safety: Generating tasks run asynchronously via task_id; credits are secured upon submission and automatically refunded if a job fails.

Hailuo 2.3 Pro Text to Video API frequently asked questions

What is the Hailuo 2.3 Pro Text to Video API?

Hailuo 2.3 Pro Text to Video is MiniMax's premier text-to-video generation model, developed for high-end cinematic motion and broadcast-grade visual production. It converts detailed natural language prompts into native 1080p full HD video clips with advanced physical simulation, volumetric lighting, and realistic camera choreography across focused 6-second sequences. Leveraging MiniMax's high-capacity multi-modal diffusion architecture, it renders realistic fluid, cloth, and character dynamics while maintaining pristine edge definition and frame-to-frame coherence. Developers can integrate the model programmatically through Vidgo's REST API or test creative ideas immediately in the interactive web playground above.

Does Hailuo 2.3 Pro Text to Video output native 1080p resolution?

Yes, Hailuo 2.3 Pro Text to Video outputs native 1080p full HD resolution directly from the generative model rather than relying on spatial upscaling, delivering razor-sharp textures, authentic depth of field, and filmic detail.

Why is Hailuo 2.3 Pro Text to Video limited to a 6-second duration?

The Pro tier is architected to maximize per-frame visual complexity, photoreal physics, and cinematic rendering density within a standardized 6-second window. The API only accepts duration=6; configuring any other duration value will result in a validation error.

How does pricing work for Hailuo 2.3 Pro Text to Video?

Hailuo 2.3 Pro Text to Video is billed at a fixed rate of 60 credits ($0.300 based on $0.005 per credit) per 6-second generation. Using the prompt optimizer does not incur any additional charges.

How does Hailuo 2.3 Pro Text to Video handle complex physical motion?

The model incorporates a dedicated world-simulation physical prior that accurately accounts for gravitational pull, momentum, fluid splashing, smoke dissipation, and collision dynamics, preventing unnatural morphing during rapid movements.

Can I use up to 5,000 characters in Hailuo 2.3 Pro Text to Video prompts?

Yes, the prompt field supports up to 5,000 Unicode characters, allowing directors and prompts engineers to articulate comprehensive shot lists, lighting temperatures, actor staging, lens types, and sequential narrative progression.

When should I choose Hailuo 2.3 Pro Text to Video instead of the Standard endpoint?

Choose Hailuo 2.3 Pro Text to Video when your production requires native 1080p full HD clarity, broadcast-quality lighting, and intense physical action within a 6-second timeframe. If your workflow prioritizes rapid drafting, lower unit cost (35 credits for 6s), or requires 10-second extended shots (70 credits), the Standard Text to Video endpoint is the recommended choice.