Happy Horse 1.1 Image to Video API

alibaba/happyhorse-1.1/image-to-video

Happy Horse 1.1 Image to Video animates a single still photo into a 3–15 second cinematic video clip at 720p or 1080p, with native synchronized audio, natural physical motion, and rich textural detail. It preserves the original subject identity, lighting, and composition while introducing realistic camera moves, character performance, and environment sounds.

Input
588/2500
Output
Idle

Your generated video will appear here

Add your prompt and required media, review the settings, then click Run.

1080p · 5 sec × 28 credits/sec = 140 credits ($0.700)
Continue with

Examples

Animate this first frame as one continuous five-second intimate documentary shot. Preserve this adult ceramic artist, face, clay vessel, hands, clothing and sunlit pottery studio. The wheel rotates gently while both wet hands steadily guide the rim without changing the vessel into another object. The artist glances up and calmly says exactly in Mandarin Chinese, "慢一点,形就稳了。" Clearly synchronized lips, natural small hand movements, soft wheel hum and wet-clay rubbing sounds. Very slow camera push-in, no cuts, no music. No logos, brands, advertising, captions, subtitles or watermarks.

Animate the supplied first frame in one continuous three-second underwater macro shot. Keep the same single translucent jellyfish and dark blue water. Its bell contracts once and relaxes, carrying it gently upward while fine tentacles follow smoothly and tiny suspended particles drift independently. Preserve the intricate transparent tissue and soft refracted highlights. Almost stationary camera, subtle underwater ambience, no speech or music. No logos, brands, advertising, captions, subtitles or watermarks.

Animate this exact first frame for three seconds. Preserve the layout of the old sunlit laundry room. Colored garments tumble slowly inside the single circular washing-machine window while a hanging white cotton sheet in the foreground billows gently once in a cross-breeze. Bright afternoon sunlight passes through the cloth and moves the soft shadow across the tiled floor. Fixed camera, realistic fabric folds and rotating laundry, quiet mechanical rumble and soft cloth rustle, no people, speech or music. No logos, brands, advertising, captions, subtitles or watermarks.

Happy Horse 1.1 Image to Video

Happy Horse 1.1 Image to Video brings still images to life as fully voiced, dynamic video clips. Provide exactly one initial frame image via public URL, Data URI, or Base64, then optionally add up to 2,500 characters of natural-language direction for action, camera travel, or spoken dialogue. The model maintains subject identity and original lighting from the starting image while synthesizing fluid motion and matching environmental audio in a single pass.

Key Capabilities & Advantages

  • Starting-frame identity preservationLocks character features, clothing textures, and product silhouettes from the input image throughout the animated take.

  • Automatic image aspect ratio inheritanceAdopts the natural width and height proportions of your source image, avoiding unwanted border letterboxing or forced stretching.

  • Joint acoustic synthesis from still artCreates fitting acoustic ambiance, Foley effects, and musical elements that correspond directly with visual dynamics.

  • Zero-prompt autonomous animationIntelligently analyzes visual cues to generate plausible physical motion and atmosphere even when no text prompt is supplied.

  • Director-grade camera movementAccepts instructions for steady push-ins, tracking pans, and rotational perspective shifts without distorting foreground subjects.

  • Seamless 7-language dialogue animationAnimate portraits into speaking avatars with mouth shapes and spoken lines matching English, Chinese, Japanese, Korean, German, or French.

Parameters

ParameterRequirementDescription
promptOptional

Up to 2500 Unicode characters after trimming surrounding whitespace. Optional; null, empty, and whitespace-only prompts are omitted.

image_urlsRequired

Exactly one first-frame image: public HTTP(S) URL, image Data URI, or raw Base64.

resolutionOptional

720p or 1080p. Default: 1080p.

Default1080p
durationOptional

Integer from 3 to 15 seconds. Default: 5.

Default5
seedOptional

Optional integer from 0 to 2147483647. Omitted when not specified.

enable_safety_checkerOptional

Optional boolean. Omitted when not specified.

How to Call Happy Horse 1.1 Image to Video API

  1. Select a high-quality source imageChoose a still photo with sharp focus, balanced illumination, and distinct subject boundaries.

  2. Provide image inputUpload your image or specify a public image URL in the image_urls array parameter.

  3. Add optional motion and audio guidanceOptionally write a prompt specifying camera trajectory, character gesture, or spoken dialogue lines.

  4. Configure clip length and resolutionSelect an output duration between 3 and 15 seconds and pick 720p or 1080p resolution.

  5. Dispatch the generation requestSubmit via the API endpoint or press generate in the online console.

  6. Download the animated videoPoll the status endpoint with task_id until marked finished, then retrieve the MP4 file.

Pricing

Cost = output seconds × resolution rate. 1 credit = $0.005; all three modes use the same rates. Failed tasks are refunded automatically.

UsageRateDetails
720p22 credits/s ($0.11/s)5 seconds: 110 credits ($0.55)
1080p28 credits/s ($0.14/s)5 seconds: 140 credits ($0.70)

Best Use Cases

  • E-commerce product animationTurn static product photography into rotating, dynamic video advertisements with studio lighting gleams and gentle ambient audio.

  • Portrait and avatar animationTransform character stills or corporate headshots into talking spokesperson clips complete with multilingual lip-sync.

  • Historical photo restorationBring vintage photographs and historical portraits to life with authentic ambient soundscapes and gentle natural motion.

  • Concept art and matte painting motionAnimate static landscape illustrations into breathing cinematic vistas with wind rustle, flowing water, and camera sweeps.

Pro Tips

  • When animating people, describe micro-expressions (such as a subtle smile, blinking, or nodding) alongside main camera movement to keep facial acting lifelike.
  • If you desire speech, include quoted dialogue lines and specify the language in your prompt (e.g. 'Character looks toward lens and says warmly in Japanese: Konnichiwa').
  • Leave prompt empty when you want the AI to naturally interpret outdoor landscapes, flowing rivers, or atmospheric clouds based purely on visual clues.
  • For product showcase animations, prescribe slow orbit or push-in camera tracks ('slow camera push-in highlighting product reflections on metallic rim') for commercial appeal.

Usage Notes

  • Pass exactly one image URL or data string in image_urls; multi-image reference workflows belong to Reference to Video.
  • Do not submit aspect_ratio; the output video automatically conforms to the aspect ratio of your uploaded image.
  • Audio is synthesized jointly with animation; silent generation can be requested by specifying ambient room tone in your prompt.
  • Prompts are optional (up to 2,500 characters); omit or provide natural-language motion direction as desired.

Related Models

Happy Horse 1.1 Image to Video API frequently asked questions

What is the Happy Horse 1.1 Image to Video API?

Happy Horse 1.1 Image to Video is an Alibaba model for video generation from single still images. It animates a starting image into 3–15 second cinematic video clips at 720p or 1080p with native synchronized audio, natural physical motion, and rich textural detail. Built on Alibaba's unified single-stream self-attention Transformer architecture, it preserves the source image's character identity, lighting, and composition while introducing realistic camera moves and sound. You can call it programmatically or try it from the playground above.

How does Happy Horse 1.1 Image to Video preserve character identity from a photo?

The model anchors visual feature representations directly to the input frame, ensuring that facial geometry, hairstyles, distinctive attire, and ambient lighting remain consistent throughout dynamic camera movements and gestures.

Can Happy Horse 1.1 Image to Video generate video without a text prompt?

Yes. When no prompt is provided or when the prompt is empty, the model autonomously analyzes scene elements, predicting realistic natural motion such as wind blowing, flowing water, or organic character breathing alongside matching environmental sound.

Does Happy Horse 1.1 Image to Video generate synchronized sound for animated photos?

Yes. Audio synthesis is integrated into the model's single forward pass. Whether synthesizing ambient sound from image context or generating spoken dialogue from prompt instructions, sound is generated natively in sync with video frames.

How is the video aspect ratio determined in Happy Horse 1.1 Image to Video?

The resulting video automatically inherits the aspect ratio and frame orientation of the submitted input image. You do not need to pass an aspect_ratio parameter.

What image formats and resolutions work best with Happy Horse 1.1 Image to Video?

Provide clear JPG, PNG, or WEBP images with balanced lighting and sharp focus. Standard portrait or landscape images with clean subject separation yield the most stable motion and detail retention.

How do I direct camera motion in Happy Horse 1.1 Image to Video?

Specify camera techniques in your text prompt using standard cinematography terminology, such as slow zoom-in, pan right following subject, or handheld tracking shot. Keep camera instructions in a distinct sentence from subject action descriptions.

How are credits billed for Happy Horse 1.1 Image to Video?

Charges equal requested duration in seconds multiplied by the resolution rate: 22 credits per second for 720p, or 28 credits per second for 1080p. Failed generation tasks are fully refunded.