Gemini Omni Flash Reference to Video API

google/gemini-omni-flash/reference-to-video

Gemini Omni Flash Reference to Video builds 4–10 second clips from exactly three reference images plus a text prompt, with locked character consistency, native lip-synced audio, and resolution from 720p to 4K. It carries face, wardrobe, and style cues across the new scene while adding directed motion, camera language, and synchronized sound.

Input
441/20000
0/3 · remaining 3
Output
Idle

Your generated video will appear here

Add your prompt and required media, review the settings, then click Run.

720p · 6 sec = $0.750 (150 credits)
Continue with

Examples

Use the reference images only for this astronaut's face, short red hair, and white-and-orange flight suit. New scene: she steps into a vast alien crystal cave, giant luminous crystals refracting teal and violet light, she turns her head looking around in wonder. Camera follows behind her in one steady tracking shot. Audio: echoing footsteps and a low crystalline hum. Keep her identity consistent. No logos, no readable text, no watermark.

Use the reference images only for this border collie's markings, blue eye, and identity. New scene: the dog herds a small flock of sheep across a misty hillside, trotting and turning to guide them, fog drifting over green grass. Documentary-style steady wide shot. Audio: soft wind, distant sheep bells. Keep the dog's appearance consistent. No logos, no readable text, no watermark.

Use the reference images only for this white ceramic fox figurine with gold accents, its shapes and colors. New scene: the figurine comes alive in a snowy pine forest, it shakes snow off its back, blinks, and steps carefully between the trees, gentle snowfall. Magical realism, steady shot. Audio: soft snow ambience and tiny footsteps. Keep its design consistent. No logos, no readable text, no watermark.

Gemini Omni Flash Reference to Video

Gemini Omni Flash Reference to Video is Google DeepMind’s multimodal model for multi-image consistent video generation. Provide exactly three public image URLs—for example face, wardrobe, and product or style—and a cinematic prompt to output 4–10 second clips with lip-synced audio. Choose 720p, 1080p, or 4K with 16:9 or 9:16 framing—ideal for recurring characters, campaign talent continuity, product line consistency, and branded virtual-host series.

Why Choose This?

  • Three-image character lockUse exactly three references to anchor face, wardrobe, and style so recurring talent stays recognizable across shots.

  • Strong multi-subject consistencyName people and products in the prompt while the references supply visual ground truth for appearance continuity.

  • Native audio with lip syncGenerate dialogue and ambience with the picture for talking characters that match the locked look.

  • Cinematic staging on locked looksApply push-in, dolly, tracking, and lighting cues without rebuilding identity from text alone.

  • 720p to 4K delivery tiersIterate at 720p or 1080p, then step to 4K when campaign finals need finer texture and lighting detail.

  • Flexible duration and framingPick 4–10 seconds and 16:9 or 9:16 for hooks, mid-length demos, and vertical social series.

Parameters

ParameterRequirementDescription
promptRequired

String. Scene, action, camera, lighting, and audio cues; 1–20,000 characters after trimming.

image_urlsRequired

Array with exactly 3 public HTTP(S) image URLs for character, wardrobe, or style reference.

durationOptional

Integer. Output length in seconds; the Playground preselects 6.

Default64810
resolutionOptional

String. Output clarity tier; the Playground preselects 720p.

Default720p1080p4k
aspect_ratioOptional

String. Output framing; the Playground preselects 16:9.

Default16:99:16

How to Use

  1. Prepare three reference imagesGather clear public URLs for face, wardrobe, and product or style so identity cues cover appearance and branding.

  2. Write the scene promptDescribe action, camera, lighting, and audio, and name which reference drives face, outfit, or product look.

  3. Set output durationChoose 4, 6, 8, or 10 seconds (default 6) to match the narrative beat.

  4. Select resolutionPick 720p for iteration, 1080p for clearer delivery, or 4K for high-detail finals.

  5. Choose aspect ratioSelect 16:9 or 9:16 to match landscape storytelling or vertical social series.

  6. Review the cost and runCheck the cost shown on the Run button, finish references and prompt, then click Run.

  7. Preview and download the videoWhen the task finishes, preview picture and synced audio, then select Download video to save the result.

Pricing

Billed per generation by duration and resolution tier, with native audio included. 1 credit = $0.005.

UsageRateDetails
720p / 1080p4s=120, 6s=150, 8s=200, 10s=220 creditsDefault 720p / 6s costs 150 credits ($0.75).
4k4s=250, 6s=300, 8s=350, 10s=450 credits4k / 6s costs 300 credits ($1.50).

Best Use Cases

  • Recurring character seriesKeep the same talent look across episodic social shorts and campaign chapters.

  • Product line consistencyLock packaging and hero product appearance while staging new environments and camera moves.

  • Brand virtual-host contentCombine face and wardrobe references with dialogue cues for ongoing host performances.

  • Multi-shot campaign storyboardsPrototype story beats that must share one visual identity before full production.

Pro Tips

  • Assign roles to the three images in the prompt—face reference, wardrobe reference, product or style reference.
  • Keep reference lighting consistent and faces unobstructed so identity signals stay strong.
  • Reuse the same appearance wording across iterations when building a series of related clips.
  • Specify push-in, dolly, or tracking moves separately from character action to control staging.
  • Add Audio lines for dialogue language and ambience so lip sync matches the locked performer.

Usage notes

  • Gemini Omni Flash Reference to Video requires prompt plus image_urls with exactly three public image URLs.
  • Describe speech or ambience in the prompt; native audio with lip sync is included in the result.
  • Duration, resolution, and aspect_ratio configure output length, clarity, and framing.
  • After an API submission, save the returned task_id to query progress and retrieve the final media URL.

Related Models

Gemini Omni Flash Reference to Video API frequently asked questions

What is the Gemini Omni Flash Reference to Video API?

Gemini Omni Flash Reference to Video is a Google DeepMind multimodal model for generating video from multiple reference images. It creates 4–10 second clips from exactly three reference stills plus a text prompt, with locked character consistency, native lip-synced audio, and 720p to 4K output. Built on Gemini’s unified multimodal architecture, it carries face, wardrobe, and style cues into a new scene while adding directed motion and synchronized sound. You can call it programmatically or try it from the playground above.

How many reference images does Gemini Omni Flash Reference to Video require?

Submit exactly three public image URLs in image_urls. Typical setups pair face, wardrobe, and product or style references; naming each role in the prompt strengthens consistency.

How does Gemini Omni Flash Reference to Video keep character consistency?

The three stills supply visual ground truth for appearance, while the prompt restates face, outfit, and branding details. Reuse the same reference set and appearance wording when generating related clips in a series.

Does Gemini Omni Flash Reference to Video support native lip sync?

Yes. Dialogue and ambience are generated with the picture and embedded in the MP4. Add language and tone cues so the locked performer speaks in sync with mouth motion.

When should I choose Gemini Omni Flash Reference to Video over Image to Video?

Use Reference to Video when one still is not enough and you need multi-image locks for face, wardrobe, and style. Use Image to Video when a single key visual already carries the full look you want to animate.

Can Gemini Omni Flash Reference to Video output 4K?

Yes. Choose 4K for campaign finals that need sharper texture and lighting. Draft at 720p or 1080p first to validate consistency, then promote selected takes.

How are Gemini Omni Flash Reference to Video credits calculated?

Credits are charged per generation by duration and resolution. At 720p/1080p, 6 seconds costs 150 credits ($0.75); at 4K the same length costs 300 credits ($1.50). See the Pricing section for full rates.