Happy Horse 1.1 Text to Video API

alibaba/happyhorse-1.1/text-to-video

Happy Horse 1.1 Text to Video transforms written descriptions into 3–15 second cinematic video clips at 720p or 1080p, with native synchronized audio, multilingual lip-sync across seven languages, and expressive physical motion. It maintains temporal coherence and shot staging across complex scene prompts while generating speech, sound effects, and music in a single forward pass.

Input
654/2500
Output
Idle

Your generated video will appear here

Add your prompt and required media, review the settings, then click Run.

1080p · 5 sec × 28 credits/sec = 140 credits ($0.700)
Continue with

Examples

One continuous cinematic medium shot of an adult female glaciologist in a mustard expedition jacket on a safe rocky overlook beside a blue glacier in clear cold daylight. Her uncovered face remains large and readable. Over five seconds she turns slightly toward the camera, quietly says exactly in English, "Listen—the ice is moving.", then listens as a distant deep ice crack rolls across the valley. Slow subtle camera push-in, realistic breath and restrained expression. Accurate visible lip synchronization, soft wind and one distant cracking sound, no music. No collapse or disaster. No logos, brands, advertising, captions, subtitles or watermarks.

Three-second continuous side-tracking sports documentary shot at a sunlit concrete skatepark. One adult BMX rider wearing a matte teal helmet performs a single low bunny hop over a painted line and lands both wheels smoothly. Begin already rolling, show the complete small hop and finish rolling away. Plausible bicycle geometry, two wheels stay attached, natural body balance, crisp tire contact and short landing thud, no speech or music. No logos, brands, advertising, captions, subtitles or watermarks.

A three-second handmade stop-motion scene in a miniature rehearsal room, square composition. A charming human-shaped coral-clay drummer wearing a turquoise shirt sits at one small yellow snare drum and practices a short relaxed rhythm using two wooden sticks. Each audible dry drum tap follows a visible stick contact, with small natural head nods and body bounces. The left and right hands remain anatomically coherent and the drum stays in place. Visible clay fingerprints, tactile felt acoustic panels, warm rehearsal-room light, steady frontal camera, consistent character and instrument. Only the natural snare taps and quiet room ambience, no speech or background music. No logos, brands, advertising, captions, subtitles or watermarks.

Happy Horse 1.1 Text to Video

Happy Horse 1.1 Text to Video generates complete video clips with native synchronized sound from text prompts alone. Describe the environment, characters, camera direction, and auditory cues in up to 2,500 characters, then choose an integer duration from 3 to 15 seconds, a resolution of 720p or 1080p, and your preferred framing from 9 aspect ratios. The model produces video frames, speech dialogue, ambient Foley, and background music together in one unified step.

Key Capabilities & Advantages

  • Joint audio-video generationProduces visual frames and matching audio tracks simultaneously, eliminating secondary dubbing or separate audio post-processing pipelines.

  • Multilingual native lip-syncAligns character mouth movements to spoken dialogue accurately in English, Mandarin, Cantonese, Japanese, Korean, German, and French.

  • Dynamic and grounded motionDelivers smooth physical movement in fast-action sequences like athletics, dance, and chases with significantly reduced stutter.

  • Storyboard-style prompt timingUnderstands time markers such as 0-4s and 4-8s to choreograph sequential actions, shot transitions, and spoken lines across the clip.

  • Flexible cinematic aspect ratiosSupports 9 native framing options including widescreen 21:9, standard 16:9, square 1:1, and vertical 9:16 for direct multi-platform delivery.

  • Predictable second-based billingComputes cost directly from requested duration and resolution, ensuring transparent budget estimation before running.

Parameters

ParameterRequirementDescription
promptRequired

Up to 2500 Unicode characters after trimming surrounding whitespace. A nonblank prompt is required.

resolutionOptional

720p or 1080p. Default: 1080p.

Default1080p
durationOptional

Integer from 3 to 15 seconds. Default: 5.

Default5
aspect_ratioOptional

Supported values: 21:9, 16:9, 4:3, 1:1, 3:4, 4:5, 5:4, 9:16, 9:21. Default: 16:9.

Default16:9
seedOptional

Optional integer from 0 to 2147483647. Omitted when not specified.

enable_safety_checkerOptional

Optional boolean. Omitted when not specified.

How to Call Happy Horse 1.1 Text to Video API

  1. Establish the subject and settingBegin your prompt by defining the environment, lighting style, and primary character or focal object in clear detail.

  2. Describe motion and pacingOutline how characters move, specify camera behavior such as dolly shots or tracking pans, and define timing across the clip.

  3. Specify dialogue and acoustic cuesInclude dialogue lines in quotation marks, mention the spoken language for lip-sync, and describe ambient sound or background score.

  4. Select duration, resolution, and ratioChoose an output length from 3 to 15 seconds, pick 720p or 1080p, and configure your target aspect ratio from the 9 supported options.

  5. Submit generation taskCall the asynchronous API endpoint or click generate in the playground console with your configured parameters.

  6. Retrieve final videoQuery the task status using the returned task_id and download the resulting MP4 video with embedded audio upon completion.

Pricing

Cost = output seconds × resolution rate. 1 credit = $0.005; all three modes use the same rates. Failed tasks are refunded automatically.

UsageRateDetails
720p22 credits/s ($0.11/s)5 seconds: 110 credits ($0.55)
1080p28 credits/s ($0.14/s)5 seconds: 140 credits ($0.70)

Best Use Cases

  • Commercial concept visualizationRapidly turn written advertising scripts into dynamic video concepts complete with voiceover and background music for client pitches.

  • Social media short-form contentProduce engaging vertical 9:16 narrative clips with spoken punchlines and synchronized reactions tailored for TikTok, Shorts, and Reels.

  • Game cinematics and cutscenesGenerate dramatic character performances and action vignettes with tailored environmental soundscapes for game development previsualization.

  • Multilingual ad localizationWrite prompts featuring local dialogue in French, German, Japanese, Korean, Cantonese, or Mandarin for global marketing campaigns.

Pro Tips

  • Use timestamped dialogue blocks (for example, '0-3s: Character speaks in English; 3-5s: Camera dollies back with ambient room tone') to direct pacing cleanly.
  • Explicitly name sound details in your prompt; mentioning footstep textures, rain ambience, or orchestral tone helps the joint audio engine construct richer soundscapes.
  • Separate character action instructions from camera movement instructions into distinct sentences for tighter visual adherence.
  • Test short 5-second 720p generations during creative ideation before ordering 15-second 1080p deliverables.

Usage Notes

  • Prompts must contain between 1 and 2,500 Unicode characters; whitespace-only submissions are rejected before deduction.
  • Generation is asynchronous; query status periodically with task_id until the task transitions to finished or failed.
  • Audio tracks are generated natively within the video file; no auxiliary audio stream or external dubbing tool is required.
  • Aspect ratio defaults to 16:9; choose from 9 native ratios to match your presentation target without letterboxing.

Related Models

Happy Horse 1.1 Text to Video API frequently asked questions

What is the Happy Horse 1.1 Text to Video API?

Happy Horse 1.1 Text to Video is an Alibaba model for video generation from text prompts. It generates 3–15 second cinematic video clips at 720p or 1080p with native synchronized audio, multilingual lip-sync across seven languages, and expressive physical motion. Built on Alibaba's unified single-stream self-attention Transformer architecture, it preserves coherent temporal continuity and lighting realism while generating speech, sound effects, and music in a single forward pass. You can call it programmatically or try it from the playground above.

Can Happy Horse 1.1 Text to Video generate dialogue with lip-sync from text?

Yes. When dialogue lines are included in quotation marks within your prompt, the model synchronizes character facial musculature and mouth shapes to the spoken phonemes while rendering matching vocal audio directly in the video file.

Which languages support lip-sync in Happy Horse 1.1 Text to Video?

Happy Horse 1.1 Text to Video supports high-accuracy lip-sync in English, Mandarin, Cantonese, Japanese, Korean, German, and French. Specify the language alongside dialogue text in your prompt to guide pronunciation and acoustic delivery.

How do I structure multi-scene timing in Happy Horse 1.1 Text to Video prompts?

Use clear second markers such as 0-4s and 4-8s in your text prompt. The model reads these chronological ranges to transition camera positions, sequence character actions, and pace spoken phrases across your chosen duration.

What durations and resolutions does Happy Horse 1.1 Text to Video offer?

You can choose an integer duration from 3 to 15 seconds, with 5 seconds set as default. Resolutions include 720p for fast drafting and 1080p for final production deliverables.

Which aspect ratios are available for Happy Horse 1.1 Text to Video?

The endpoint provides 9 aspect ratios: 21:9, 16:9, 4:3, 1:1, 3:4, 4:5, 5:4, 9:16, and 9:21. Specify your desired ratio in the aspect_ratio field to match your intended presentation format without cropping.

How does audio generation work in Happy Horse 1.1 Text to Video?

Audio and visual frames are synthesized jointly during inference. Describing acoustic elements such as footstep materials, weather ambience, crowd murmur, and musical mood enables the model to produce coordinated sound effects alongside character speech.

How are credits calculated for Happy Horse 1.1 Text to Video?

Billing is computed by multiplying the requested duration in seconds by the resolution rate: 22 credits per second for 720p, and 28 credits per second for 1080p. Credits are deducted upon submission and refunded automatically if generation cannot complete.