Gemini Omni 1.1 Flash Image-to-Video API

google/gemini-omni-1.1-flash/image-to-video

Gemini Omni 1.1 Flash (Image-to-Video) turns a start image and text prompt into video with native audio, optional end-frame guidance, and 360p to 4K output. Build action, camera movement, and sound around the subject and composition of your opening image, adding an end frame to guide the closing shot.

Input
475/20000
Output
Idle

Your generated video will appear here

Add your prompt and required media, review the settings, then click Run.

720p · 8 sec = $0.375 (75 credits)
Continue with

Examples

Begin exactly on the first frame and finish on the last frame. One continuous 8-second locked close-up: the folded paper crane's wings slowly uncrease and lift a few centimeters until they match the end still. Preserve the same crane, paper color, table, window light, and camera. Native audio: dry paper flex, a faint wooden-table creak, quiet room tone. No logos, no readable text, no products, no packaging, no prices, no CTA, no advertising, no watermark, no brand marks.

Begin from the first frame. One continuous 4-second locked wide shot: a sudden countryside gust lifts the cream sheet and the faded shirt, they billow and slap the line, then settle. Preserve the same posts, yard, dusk light, and camera. Native audio: fabric slap, dry grass wind, distant insects. Household laundry only, not a clothing ad. No logos, no readable text, no products, no packaging, no prices, no CTA, no advertising, no watermark, no brand marks.

Begin on the first frame and finish on the last. One continuous 4-second 9:16 shot: the same man slowly lifts his chin from the monstera leaf and looks toward camera, keeping glasses, indigo shirt, face, and greenhouse layout. Locked camera, same left-wall morning light. Native audio: leaf rustle, a drip from irrigation, soft greenhouse ambience. No logos, no readable text, no products, no packaging, no prices, no CTA, no advertising, no watermark, no brand marks.

Gemini Omni 1.1 Flash Image-to-Video

Gemini Omni 1.1 Flash Image-to-Video turns a start image and text prompt into a moving clip with native audio, with an optional end image to guide the closing composition. Use your chosen image as the visual starting point for character action, product shots, and camera movement that bring still assets into video.

Why Choose This?

  • Animation from a start frameBegin with the subject, lighting, and composition of a chosen image to turn product artwork, portraits, or scene stills into moving footage.

  • End-frame guidanceAdd an optional end frame to define a target pose, product position, or closing composition and give the clip a clear visual destination.

  • Transitions between framesProvide start and end images, then describe the action and camera movement between them to plan composition changes and product transitions.

  • Prompt-directed motionDescribe a head turn, moving fabric, a camera push-in, or an orbit to build an existing image around a specific action.

  • Native sound creationAdd footsteps, ambient sound, or music to the prompt to design a soundtrack around the visible action.

  • Flexible output specificationsCombine 360p to 4K resolution, landscape or portrait framing, and a 4–10 second duration for product presentations and social content.

Parameters

ParameterRequirementDescription
promptRequired

String. Directs action, camera movement, visual change, and sound after the start frame; 1–20,000 characters after trimming.

image_urlsRequired

String array with one or two public URLs. The first image is the start frame; the optional second image is the end frame. JPEG, PNG, or WebP, up to 30 MiB each.

durationOptional

Integer. Sets output length; the Playground preselects 8 seconds.

Default84610
resolutionOptional

String. Sets output resolution; the Playground preselects 720p.

Default720p360p1080p4k
aspect_ratioOptional

String. Controls output framing; the Playground preselects 16:9.

Default16:99:16

How to Use

  1. Upload a start imageChoose a JPEG, PNG, or WebP image with a clear subject and composition, up to 30 MiB, as the opening frame.

  2. Add an optional end frameTo guide the closing pose or composition, add an end image after the start frame using the same format and size requirements.

  3. Describe action and soundDescribe the subject action, camera movement, lighting, and audio after the opening frame; with an end image, explain how the shot reaches it.

  4. Choose a resolutionChoose 360p, 720p, 1080p, or 4K, with 720p selected by default, to match the output specifications for editing or presentation.

  5. Choose a durationChoose 4, 6, 8, or 10 seconds, with 8 seconds selected by default, to plan the clip around its main action and pacing.

  6. Choose an aspect ratioChoose 16:9 landscape or 9:16 portrait, with 16:9 selected by default, and frame the subject for the intended layout.

  7. Review the cost and runReview the cost shown on the Run button, complete the required uploads and prompt, then click Run.

  8. Preview and download the videoWhen the task finishes, preview the video and sound in the output panel, then select Download video to save the result.

Pricing

Billed per generation based on video duration and resolution tier, with native audio included in the result. 1 credit = $0.005.

UsageRateDetails
360p / 720p / 1080p4s=45, 6s=60, 8s=75, 10s=90 creditsDefault 720p / 8s costs 75 credits ($0.375).
4k4s=105, 6s=120, 8s=135, 10s=150 credits4k / 8s costs 135 credits ($0.675).

Best Use Cases

  • Animated product visualsUse a product still as the start frame and describe camera movement and lighting changes to create footage for brand presentations.

  • Portrait animationStart with a portrait and describe a head turn, expression, or clothing movement for character presentations and story shots.

  • Start-to-end transitionsProvide opening and closing images and describe the action between them to plan product transformations or scene transitions.

  • Storyboard animationUse a selected storyboard image as the start frame, then add camera and sound direction for a shot preview shared across creative teams.

Pro Tips

  • Choose a start image with a clear subject outline, lighting, and composition to establish the visual starting point.
  • Describe action that continues from the image, such as “The person raises their cup as the camera slowly moves closer,” to define what happens next.
  • When adding an end frame, choose consistent subject features and visual styling, then describe the action leading into the final pose.
  • Describe the facial features, clothing, or setting to carry forward in one sentence, then specify camera motion separately to distinguish retained details from movement.
  • Describe sounds tied to visible events, such as footsteps or a product opening, to connect the soundtrack to the action.

Usage notes

  • Gemini Omni 1.1 Flash Image-to-Video generates video from a start image and text prompt, using prompt and image_urls as its main inputs. An optional second image guides the end frame, while duration, resolution, and aspect ratio configure the output.
  • Describe speech, music, or ambient sound in the prompt; native audio is included in the result.
  • Use publicly accessible HTTP(S) URLs for API media inputs so the service can retrieve the files.
  • Save the task_id returned by an API submission to query progress and retrieve the result.

Gemini Omni 1.1 Flash Image-to-Video API frequently asked questions

What is the Gemini Omni 1.1 Flash Image-to-Video API?

Gemini Omni 1.1 Flash Image-to-Video is a Google model for generating video from images. It animates static starting frames into clips up to 4K resolution with native synchronized audio based on text instructions, and supports uploading an optional closing frame for end-state control. Built on Gemini's multimodal intelligence architecture, it faithfully preserves the source subject identity, clothing textures, and lighting while introducing natural physical dynamics. You can call it programmatically or try it from the playground above.

Does Gemini Omni 1.1 Flash Image-to-Video support first and last frame interpolation?

Yes. When you provide a second image URL in the image_urls array as an end frame, the model computes spatial transitions and structural displacements between both stills, producing a seamless continuous camera take from the opening frame to the closing frame.

How does Gemini Omni 1.1 Flash Image-to-Video preserve subject identity?

The model anchors facial likeness, outfit details, and ambient lighting directly from the starting image. Providing high-resolution assets with clear subject boundaries and describing specific motion trajectories in the prompt provides unambiguous guidance for consistent subject animation.

Can Gemini Omni 1.1 Flash Image-to-Video synthesize ambient audio for static images?

Yes. With native audio-visual cross-modal reasoning, the model infers the visual context (such as an ocean beach, a bustling cafe, or a rainy street) from the start frame and automatically synthesizes matching environmental audio and physical foley without requiring manual sound uploads.

How can I create seamless looping clips with Gemini Omni 1.1 Flash Image-to-Video?

To produce a seamless loop, pass the exact same image as both the first and last frame in image_urls, and specify a subtle cyclic movement or steady circular camera orbit in your prompt. The model completes the action cycle within the set duration and returns cleanly to the initial composition.

What image formats are recommended for Gemini Omni 1.1 Flash Image-to-Video?

The endpoint accepts JPEG, PNG, and WebP images up to 20MB. Clear exposure and well-defined contours give the model the richest visual detail to extract facial geometry, material textures, and spatial depth accurately.

Can Gemini Omni 1.1 Flash Image-to-Video animate an image without a text prompt?

Yes. If no prompt is provided, the model infers plausible real-world physics from the image composition to generate organic micro-movements, such as hair sway, fabric drift, or water ripples. Adding text instructions allows you to guide explicit camera moves or deliberate character choreography.