Happy Horse 1.1 Reference to Video API

alibaba/happyhorse-1.1/reference-to-video

Happy Horse 1.1 Reference to Video turns 1 to 9 reference images and a guiding text prompt into a 3–15 second cinematic video clip at 720p or 1080p, with multi-subject identity consistency, native synchronized audio, and multilingual lip-sync across seven languages. It preserves distinct character features, wardrobe designs, and object silhouettes across scene changes without identity drift.

Input
694/2500
0/9 · remaining 9
Output
Idle

Your generated video will appear here

Add your prompt and required media, review the settings, then click Run.

1080p · 5 sec × 28 credits/sec = 140 credits ($0.700)
Continue with

Examples

Bind [image_1] to the little handmade felt explorer and [image_2] to the explorer’s exact copper lantern. A three-second continuous side-tracking shot inside a dark miniature cave with softly glowing pale-green fungi. The explorer takes two slow steps while holding the lantern by its handle; its amber light brushes the felt face and stone wall. Preserve the teal coat, round tan felt face, ochre cap, and the lantern’s hexagonal copper cage and blue glass star. Tactile stop-motion craftsmanship, steady coherent anatomy and prop shape, soft felt footsteps and tiny metal-handle creak, no dialogue or music. No logos, brands, advertising, captions, subtitles or watermarks.

Use [image_1] as the exact adult dancer and lilac rehearsal costume; use [image_2] as the empty theatre stage and lighting. In a full-body vertical shot on that stage, the dancer performs one slow controlled quarter-turn with arms softly opening, ending in balance. Preserve face, body proportions, lilac fabric and hairstyle throughout; feet remain grounded except for a natural heel lift. Three seconds, fixed camera, realistic fabric motion, soft shoe brush and subtle theatre room ambience, no audience, music or dialogue. No logos, brands, advertising, captions, subtitles or watermarks.

Happy Horse 1.1 Reference to Video

Happy Horse 1.1 Reference to Video solves identity drift across multi-shot narratives and complex scene staging. Submit 1 to 9 reference images representing one or more characters, garments, props, or background environments via public URLs, Data URIs, or Base64. A required prompt of up to 2,500 characters directs actions, dialogue, interactions, and camera choreography while preserving identity, rendering synchronized sound, and framing into your choice of 9 aspect ratios.

Key Capabilities & Advantages

  • Multi-image reference conditioning (1–9 images)Accepts 1 to 9 reference images to define multiple camera angles, costumes, distinct characters, or key props within a single generation task.

  • Multi-character identity anchoringMaintains separate visual identities for multiple characters across interaction scenes, preventing facial swapping or blending.

  • Flexible prompt-based character bindingBinds reference images either naturally through detailed descriptive prompts or explicitly using bracketed tokens like [image_1] and [image_2].

  • Native synchronized dialogue and FoleySynthesizes vocal lines with accurate lip movements and environment Foley simultaneously without third-party audio pipelines.

  • Multi-angle visual groundingCombines front, side, and three-quarter view reference photos to construct consistent 3D volume through complex camera orbits.

  • Nine native aspect ratio framingsProvides full flexibility with 9 supported aspect ratios including 16:9, 9:16, 21:9, and 1:1, independent of reference image dimensions.

Parameters

ParameterRequirementDescription
promptRequired

Up to 2500 Unicode characters after trimming surrounding whitespace. A nonblank prompt is required.

reference_image_urlsRequired

One to nine ordered reference images: public HTTP(S) URLs, image Data URIs, or raw Base64.

resolutionOptional

720p or 1080p. Default: 1080p.

Default1080p
durationOptional

Integer from 3 to 15 seconds. Default: 5.

Default5
aspect_ratioOptional

Supported values: 21:9, 16:9, 4:3, 1:1, 3:4, 4:5, 5:4, 9:16, 9:21. Default: 16:9.

Default16:9
seedOptional

Optional integer from 0 to 2147483647. Omitted when not specified.

enable_safety_checkerOptional

Optional boolean. Omitted when not specified.

How to Call Happy Horse 1.1 Reference to Video API

  1. Prepare 1–9 reference imagesCollect clear reference images of your subjects, props, or costumes with distinct details.

  2. Pass ordered reference arrayProvide the image URLs or data strings inside the reference_image_urls parameter array.

  3. Draft direction promptWrite a required prompt (up to 2,500 chars) detailing scene events, speaker dialogue, and character interactions.

  4. Select duration, resolution, and ratioChoose 3 to 15 seconds, pick 720p or 1080p, and select your preferred framing from 9 aspect ratios.

  5. Trigger asynchronous generationSubmit your payload to the unified endpoint or trigger via the interactive playground.

  6. Download the consistent videoQuery progress with task_id until finished, then access the rendered video via data.files[].file_url.

Pricing

Cost = output seconds × resolution rate. 1 credit = $0.005; all three modes use the same rates. Failed tasks are refunded automatically.

UsageRateDetails
720p22 credits/s ($0.11/s)5 seconds: 110 credits ($0.55)
1080p28 credits/s ($0.14/s)5 seconds: 140 credits ($0.70)

Best Use Cases

  • Multi-scene episodic video and web dramaKeep recurring actors and signature costumes looking identical across consecutive shots, dialogue scenes, and action sequences.

  • Multi-character dialogue scenesSupply separate reference portraits for two or more characters having a conversation with alternating dialogue and synchronized lip motion.

  • Consistent commercial brand assetsFeature proprietary mascots, corporate ambassadors, or packaged goods across diverse seasonal advertising campaigns.

  • Virtual fashion and product lookbooksShowcase garments and accessories on consistent models under varying lighting conditions, movement dynamics, and settings.

Pro Tips

  • When directing two characters, supply portrait photos for each and use explicit references in your prompt (e.g. 'The woman in blue [image_1] speaks in English, while the man in jacket [image_2] listens attentively') to avoid identity mix-ups.
  • Providing multiple angles of the same character (front, profile, three-quarter) allows the model to render smooth 360-degree orbit cameras without feature deformation.
  • Ensure reference images are well-illuminated and sharp; avoid heavy compression or extreme lens distortion for optimal fidelity.
  • Unlike Image to Video which inherits the first frame's ratio, Reference to Video requires an explicit aspect_ratio choice; select the format that matches your distribution channel.

Usage Notes

  • reference_image_urls accepts between 1 and 9 image URLs; providing an empty array causes task rejection before deduction.
  • A nonblank prompt is required (up to 2,500 characters) to define the action and role of the references in the scene.
  • Aspect ratio is configurable (defaults to 16:9); choose from 9 native ratios regardless of reference image dimensions.
  • Generates embedded audio with character speech, Foley, and music in the final MP4 file.

Related Models

Happy Horse 1.1 Reference to Video API frequently asked questions

What is the Happy Horse 1.1 Reference to Video API?

Happy Horse 1.1 Reference to Video is an Alibaba model for multi-reference video generation. It turns 1 to 9 reference images and a guiding text prompt into 3–15 second cinematic video clips at 720p or 1080p with multi-subject identity consistency, native synchronized audio, and multilingual lip-sync across seven languages. Built on Alibaba's unified single-stream self-attention Transformer architecture, it preserves distinct character features, wardrobe designs, and object silhouettes across scene changes without identity drift. You can call it programmatically or try it from the playground above.

How many reference images can I submit to Happy Horse 1.1 Reference to Video?

You can provide between 1 and 9 reference images in the reference_image_urls array. Use multiple images to supply different camera angles of one character, or provide distinct portraits for multiple interacting characters.

How do I reference specific images in Happy Horse 1.1 Reference to Video prompts?

You can refer to characters naturally by their visual attributes, or use explicit positional tags like [image_1] and [image_2] matching the order in reference_image_urls (for example, '[image_1] hands the package to [image_2]').

Can Happy Horse 1.1 Reference to Video animate multiple characters at once?

Yes. By providing distinct reference images for each character and specifying their respective actions and dialogue in the prompt, the model maintains separate appearances and synchronized lip motion for each speaker.

Is a text prompt required for Happy Horse 1.1 Reference to Video?

Yes. Unlike Image to Video which can run autonomously without text, Reference to Video requires a nonblank prompt (up to 2,500 characters) to instruct the model on how the referenced subjects should behave and interact.

Which aspect ratios are supported in Happy Horse 1.1 Reference to Video?

The endpoint supports 9 native aspect ratios: 21:9, 16:9, 4:3, 1:1, 3:4, 4:5, 5:4, 9:16, and 9:21. You must select your target ratio in the aspect_ratio parameter.

Does Happy Horse 1.1 Reference to Video support native lip-sync with multiple references?

Yes. Character dialogue included in quotation marks within your prompt is synthesized into natural vocal delivery and synchronized mouth shapes across any of the 7 supported languages.

How are credits billed for Happy Horse 1.1 Reference to Video?

Billing follows the standard rate based on generated video duration and resolution: 22 credits per second for 720p, and 28 credits per second for 1080p. The number of reference images submitted does not incur additional per-image charges.