Gemini Omni Flash Image to Video API

google/gemini-omni-flash/image-to-video

Gemini Omni Flash Image to Video animates one still image into a 4–10 second clip with subject preservation, native lip-synced audio, and resolution from 720p to 4K. It keeps identity, composition, and style from the source frame while adding directed motion, camera moves, and synchronized sound.

Input
346/20000
Output
Idle

Your generated video will appear here

Add your prompt and required media, review the settings, then click Run.

720p · 6 sec = $0.750 (150 credits)
Continue with

Examples

Start exactly on the input image. The lighthouse beam rotates and sweeps across the darkening sea, waves crash against the rocks sending up white spray, clouds drift slowly across the dusk sky. Steady cinematic wide shot, no cuts. Audio: ocean waves and wind. Keep the same lighthouse, cliff, and colors. No logos, no readable text, no watermark.

Start exactly on the input image. Thin mist drifts slowly across the still mountain lake, a paddle dips into the water from the canoe creating gentle expanding ripples, soft pink morning light shifts subtly. Camera locked. Audio: quiet water lapping and a distant bird. Keep the same lake, canoe, and mountains. No logos, no readable text, no watermark.

Start exactly on the input image. Dust motes drift and swirl inside the sunbeam, a gentle breeze flips the pages of the open book on the desk, the light flickers subtly. Camera locked. Audio: quiet room tone with soft page turns. Keep the same library, desk, and lighting. No logos, no readable text, no watermark.

Gemini Omni Flash Image to Video

Gemini Omni Flash Image to Video is Google DeepMind’s multimodal model for turning a single reference still into motion. Upload one public image URL, add a motion and audio prompt, then generate a 4–10 second clip with lip-synced dialogue and ambience. Choose 720p, 1080p, or 4K with 16:9 or 9:16 framing—ideal for product still animation, portrait talking clips, brand key-visual motion, and social conversions from existing creatives.

Why Choose This?

  • Single-image motion from brand stillsAnimate one product or portrait still into a watchable clip without rebuilding the scene from scratch.

  • Subject and composition preservationKeeps face, wardrobe, product shape, and framing cues from the source image while adding physically plausible motion.

  • Native audio with lip syncGenerates dialogue and ambience with the picture so talking-head and product demos ship with synced sound.

  • Prompted camera and lighting controlGuide push-in, dolly, tracking, and light mood on top of the still to stage reveals and emotional beats.

  • 720p to 4K delivery tiersDraft at 720p or 1080p, then render 4K when texture and lighting detail matter for final delivery.

  • Flexible duration and framingPick 4–10 seconds and 16:9 or 9:16 to match hooks, demos, and vertical social layouts.

Parameters

ParameterRequirementDescription
promptRequired

String. Motion, camera, lighting, and audio cues; 1–20,000 characters after trimming.

image_urlsRequired

Array with exactly 1 public HTTP(S) image URL used as the source still.

durationOptional

Integer. Output length in seconds; the Playground preselects 6.

Default64810
resolutionOptional

String. Output clarity tier; the Playground preselects 720p.

Default720p1080p4k
aspect_ratioOptional

String. Output framing; the Playground preselects 16:9.

Default16:99:16

How to Use

  1. Prepare the source stillProvide one clear public image URL showing the subject, product, or key visual you want to animate.

  2. Write the motion promptDescribe action, camera move, lighting change, and dialogue or ambience while stating what identity details to keep.

  3. Set output durationChoose 4, 6, 8, or 10 seconds (default 6) to match the beat you want from the still.

  4. Select resolutionPick 720p for iteration, 1080p for clearer delivery, or 4K for high-detail finals.

  5. Choose aspect ratioSelect 16:9 or 9:16 to match landscape storytelling or vertical social placement.

  6. Review the cost and runCheck the cost shown on the Run button, finish uploads and prompt, then click Run.

  7. Preview and download the videoWhen the task finishes, preview picture and synced audio, then select Download video to save the result.

Pricing

Billed per generation by duration and resolution tier, with native audio included. 1 credit = $0.005.

UsageRateDetails
720p / 1080p4s=120, 6s=150, 8s=200, 10s=220 creditsDefault 720p / 6s costs 150 credits ($0.75).
4k4s=250, 6s=300, 8s=350, 10s=450 credits4k / 6s costs 300 credits ($1.50).

Best Use Cases

  • Product still to motion adsTurn pack shots and hero stills into short demos with ambient sound for ecommerce and paid social.

  • Portrait talking clipsAnimate a clear face still with speech cues for virtual hosts, reactions, and spokesperson drafts.

  • Brand key-visual activationBring campaign stills to life with controlled camera moves before full video production.

  • Social creative remixConvert existing 9:16 or 16:9 creatives into fresh motion variants for A/B testing.

Pro Tips

  • Use a sharp, well-lit still with a clear subject silhouette so identity cues stay readable in motion.
  • State what to preserve—face, wardrobe, product labeling—then describe the new action separately.
  • Name camera moves such as slow push-in or side tracking to avoid random framing drift.
  • Add Audio cues for dialogue language and ambience tied to visible actions.
  • Keep one primary motion arc per clip; complex choreography works better as a second iteration.

Usage notes

  • Gemini Omni Flash Image to Video requires prompt plus image_urls with exactly one public image URL.
  • Describe speech or ambience in the prompt; native audio with lip sync is included in the result.
  • Duration, resolution, and aspect_ratio configure output length, clarity, and framing.
  • After an API submission, save the returned task_id to query progress and retrieve the final media URL.

Related Models

Gemini Omni Flash Image to Video API frequently asked questions

What is the Gemini Omni Flash Image to Video API?

Gemini Omni Flash Image to Video is a Google DeepMind multimodal model for animating a still image into video. It creates 4–10 second clips from one reference still plus a text prompt, with subject preservation, native lip-synced audio, and 720p to 4K output. Built on Gemini’s unified multimodal architecture, it keeps identity and composition from the source frame while adding directed motion and synchronized sound. You can call it programmatically or try it from the playground above.

How many images does Gemini Omni Flash Image to Video require?

Submit exactly one public image URL in image_urls. That still is the visual anchor for identity and composition; motion and sound are driven by the prompt. For multi-image consistency, use Reference to Video instead.

Does Gemini Omni Flash Image to Video preserve subject identity?

Yes. Start from a clear still and restate face, wardrobe, or product labeling in the prompt so the model anchors appearance while adding motion. Strong, uncluttered source frames improve continuity across the clip.

Does Gemini Omni Flash Image to Video generate native lip sync?

Yes. Dialogue and ambience are generated with the picture and embedded in the MP4. Add language and tone cues when you need talking performances aligned to mouth motion.

When should Gemini Omni Flash Image to Video use 4K?

Choose 4K for finals that need sharper texture and lighting detail from the still. Iterate at 720p or 1080p first—same duration options, lower credit cost—then promote selected takes.

How do I control camera motion in Gemini Omni Flash Image to Video?

Write explicit moves such as slow push-in, dolly, or side tracking, separate from subject action. Pair them with lighting cues so the still’s framing evolves intentionally.

How are Gemini Omni Flash Image to Video credits calculated?

Credits are charged per generation by duration and resolution. At 720p/1080p, 6 seconds costs 150 credits ($0.75); at 4K the same length costs 300 credits ($1.50). See the Pricing section for full rates.