Gemini Omni 1.1 Flash Reference-to-Video API

google/gemini-omni-1.1-flash/reference-to-video

Gemini Omni 1.1 Flash (Reference-to-Video) combines a text prompt with 1–7 reference images to generate video with character and style guidance, native audio, and 360p to 4K output. Bring referenced visual features into new settings, then direct the environment, action, and camera to create character stories and brand clips.

Input
657/20000
0/7 · remaining 7
Output
Idle

Your generated video will appear here

Add your prompt and required media, review the settings, then click Run.

720p · 8 sec = $0.375 (75 credits)
Continue with

Examples

Use the reference images only for this woman's face, hair, olive linen shirt, terracotta necklace, and identity. New scene, not a portrait studio and not a rainy alley: a sunny glass greenhouse at late morning. She walks two steps along a gravel path, then tips a plain metal watering can over a bed of leafy plants. Camera: medium tracking shot, eye level, one slow lateral move. Lighting: bright greenhouse sun with leaf shadows. Native audio: birds outside the glass, water from the can, footsteps on gravel. Keep identity consistent. No logos, no readable text, no products, no packaging, no prices, no CTA, no advertising, no watermark, no brand marks.

Use the references only for this exact tabby: orange-brown stripes, forehead M, white chest, white left front paw, notched right ear. New scene: a quiet afternoon wooden windowsill above a sunlit room. The cat walks slowly from left to right along the sill and pauses to look outside. Camera: locked medium side view. Lighting: warm afternoon window light. Native audio: room tone, a distant clock tick, soft paw steps on wood. Not a pet-product scene. No logos, no readable text, no products, no packaging, no prices, no CTA, no advertising, no watermark, no brand marks.

Use reference 1 only for ink-wash bamboo style, wet-paper texture, and mist. Use reference 2 only for the walking-coat silhouette. New scene: a foggy bamboo path at dawn, the silhouette walker takes three slow steps away along the path. Camera: gentle push-in, 16:9. Lighting: pale ink dawn, no neon, no rain city. Native audio: insects, soft footsteps on damp earth, distant water drip. No logos, no readable text, no products, no packaging, no prices, no CTA, no advertising, no watermark, no brand marks.

Gemini Omni 1.1 Flash Reference-to-Video

Gemini Omni 1.1 Flash Reference-to-Video combines a text prompt with 1–7 reference images to create a new video scene with native audio. Use images to guide character appearance, product features, or visual style, then describe the setting, action, and camera movement to shape character stories and branded content.

Why Choose This?

  • Character appearance referencesProvide facial features, hairstyles, and clothing through reference images to establish the visual direction for a character in a new setting.

  • Multiple reference viewsUse 1–7 images in a task, combining front, side, and detail views to describe the appearance and design of a subject.

  • Product feature referencesProvide product shape, materials, and colors through images, then describe a new setting and camera movement for a product scene.

  • Visual style continuityUse style references to communicate color, lighting, and art direction, bringing a brand’s visual language into new scene designs.

  • New scenes and actionsLet images establish the visual references while the prompt introduces a new location, action, and camera move for a character or product story.

  • Sound for the new sceneDescribe dialogue, music, and ambient sound in the same prompt to create native audio for the new scene, with landscape or portrait output.

Parameters

ParameterRequirementDescription
promptRequired

String. Assigns each reference a role and defines the new scene, camera, and sound; 1–20,000 characters after trimming.

reference_image_urlsRequired

Array of 1–7 public image URLs to guide characters, products, or visual style. JPEG, PNG, or WebP, up to 30 MiB each.

durationOptional

Integer. Sets output length; the Playground preselects 8 seconds.

Default84610
resolutionOptional

String. Sets output resolution; the Playground preselects 720p.

Default720p360p1080p4k
aspect_ratioOptional

String. Controls output framing; the Playground preselects 16:9.

Default16:99:16

How to Use

  1. Upload reference imagesAdd 1–7 JPEG, PNG, or WebP images, up to 30 MiB each, to guide the character, product, or visual style.

  2. Describe image roles and the sceneExplain what each reference contributes, then describe the target setting, subject placement, and main action.

  3. Add camera and audio directionSpecify camera movement, lighting, and speech, music, or ambient sound around the scene you want to create.

  4. Choose a resolutionChoose 360p, 720p, 1080p, or 4K, with 720p selected by default, to match the output specifications for editing or presentation.

  5. Choose a durationChoose 4, 6, 8, or 10 seconds, with 8 seconds selected by default, to plan the clip around its main action and pacing.

  6. Choose an aspect ratioChoose 16:9 landscape or 9:16 portrait, with 16:9 selected by default, and frame the subject for the intended layout.

  7. Review the cost and runReview the cost shown on the Run button, complete the required uploads and prompt, then click Run.

  8. Preview and download the videoWhen the task finishes, preview the video and sound in the output panel, then select Download video to save the result.

Pricing

Billed per generation based on video duration and resolution tier, with native audio included in the result. 1 credit = $0.005.

UsageRateDetails
360p / 720p / 1080p4s=45, 6s=60, 8s=75, 10s=90 creditsDefault 720p / 8s costs 75 credits ($0.375).
4k4s=105, 6s=120, 8s=135, 10s=150 credits4k / 8s costs 135 credits ($0.675).

Best Use Cases

  • Character story clipsUse character artwork or photos as appearance references, then describe a new location and action for introductions and story moments.

  • Product scenesUse product images as visual references and design interior, outdoor, or seasonal settings for branded video assets.

  • Brand style filmsCombine a brand mood board with scene directions to carry a chosen palette and art direction into video with sound.

  • Creative scene variationsBuild separate scenes around the same character or product references to develop several short-film directions for a brand pitch.

Pro Tips

  • Assign a purpose to each image, such as character appearance from the first, clothing from the second, and scene styling from the third.
  • Use front, side, and detail images of the same subject to show the visual features you want to reference.
  • Describe the target location, lighting, and background elements directly to establish a new space for the subject.
  • Separate appearance references from action directions, such as “Use the red coat from the reference. The person walks along a rainy street at night.”
  • Describe footsteps, ambience, or music associated with the new scene to connect the audio direction to its action and atmosphere.

Usage notes

  • Gemini Omni 1.1 Flash Reference-to-Video combines a text prompt with 1–7 reference images through prompt and reference_image_urls. Images guide characters, products, or style; duration, resolution, and aspect ratio configure the output.
  • Describe speech, music, or ambient sound in the prompt; native audio is included in the result.
  • Use publicly accessible HTTP(S) URLs for API media inputs so the service can retrieve the files.
  • Save the task_id returned by an API submission to query progress and retrieve the result.

Gemini Omni 1.1 Flash Reference-to-Video API frequently asked questions

What is the Gemini Omni 1.1 Flash Reference-to-Video API?

Gemini Omni 1.1 Flash Reference-to-Video is a Google model for generating video from multiple references. It takes 1–7 reference images alongside text prompts to reconstruct consistent character likeness, product details, or artistic aesthetics in brand-new environments with native audio. Built on Gemini's multimodal intelligence architecture, it decouples visual identity from rigid starting frames, enabling multi-angle storytelling and expressive camera direction. You can call it programmatically or try it from the playground above.

How many reference images can I upload to Gemini Omni 1.1 Flash Reference-to-Video?

You can include up to 7 reference images per request. Supplying diverse angles, such as front portraits, profile perspectives, and wardrobe or product close-ups, helps the model build a robust 3D representation that stays consistent across dynamic camera maneuvers.

Can Gemini Omni 1.1 Flash Reference-to-Video place a character into an entirely new scene?

Yes. The primary strength of this workflow is disentangling character identity or product geometry from the reference backgrounds, allowing you to transport subjects into novel environments, lighting conditions, or storylines described in your text prompt.

How do I assign roles to multiple images in Gemini Omni 1.1 Flash Reference-to-Video?

Assign distinct duties in your prompt text, such as specifying that Reference Image 1 guides the protagonist's face and hair, Reference Image 2 defines the trench coat, and Reference Image 3 provides the handheld camera prop. Explicit role tagging ensures precise feature mapping across entities.

Does Gemini Omni 1.1 Flash Reference-to-Video match sound to the target scene?

Yes. The model synthesizes native audio based on the target scene described in your prompt while taking visual cues from the references. For example, placing a character in a crowded outdoor bazaar will generate realistic crowd murmur and ambient street noise.

Will inconsistent lighting across reference photos harm Gemini Omni 1.1 Flash results?

The model features intelligent re-lighting capabilities that harmonize visual features captured under differing photographic conditions into the cohesive lighting scheme specified in your prompt. Mentioning key light directions and color temperatures further refines scene realism.

Should I choose Gemini Omni 1.1 Flash Image-to-Video or Reference-to-Video?

Choose Image-to-Video when the finished video must open exactly on the initial image's camera framing and composition. Choose Reference-to-Video when you want to carry over a character, outfit, or product identity into an entirely new opening camera setup and scene context.