Gemini Omni Flash Text to Video API

google/gemini-omni-flash/text-to-video

Gemini Omni Flash Text to Video turns text prompts into 4–10 second clips with native lip-synced audio, cinematic camera and lighting control, and resolution from 720p to 4K. It follows prompt narrative and pacing while keeping subject motion, scene continuity, and audiovisual alignment coherent across the shot.

Input
453/20000
Output
Idle

Your generated video will appear here

Add your prompt and required media, review the settings, then click Run.

720p · 6 sec = $0.750 (150 credits)
Continue with

Examples

Underwater documentary tracking shot: a green sea turtle glides through a towering kelp forest, volumetric sunbeams piercing deep blue water, a school of small silver fish scattering, fine particles drifting in the light. Camera smoothly follows the turtle at a steady distance, no cuts. Natural lighting, photorealistic. Audio: soft muffled ocean ambience and a distant whale call. No logos, no readable text, no products, no watermark, no brand marks.

View through a night train window: raindrops streak and merge on the glass while blurred city lights and neon bokeh sweep past in the dark, reflections sliding across the window frame. Moody cinematic tone, shallow depth of field. Audio: faint rain, rhythmic rail hum. No logos, no readable text, no watermark.

Extreme macro shot: a single drop of black ink falls into clear still water and blooms into slow swirling tendrils, delicate fluid dynamics, backlit against a soft white background, high-detail abstract motion, no camera movement. Silent, minimal. No logos, no readable text, no watermark.

Gemini Omni Flash Text to Video

Gemini Omni Flash Text to Video is Google DeepMind’s multimodal model for fast text-driven video creation. From a natural-language prompt alone, it outputs 4–10 second clips with synchronized speech, music, and ambient sound, plus director-style camera and lighting cues. Choose 720p, 1080p, or 4K with 16:9 or 9:16 framing—ideal for social short-form batches, ad concept drafts, brand storyboards, and creative prototyping.

Why Choose This?

  • Text-only omni video in one passDescribe subject, scene, mood, and action in natural language to generate a finished short clip with picture and sound together—no separate audio pipeline.

  • Native audio with lip syncEmbeds dialogue, music, and ambience in the MP4, aligning mouth motion and timing with spoken lines and on-screen action.

  • Cinematic camera and lighting languageResponds to push-in, dolly, tracking, orbit, and lighting cues so you can stage product reveals and story beats with clear visual direction.

  • 720p to 4K delivery tiersIterate at 720p or 1080p, then step up to 4K when you need sharper texture and lighting detail for final delivery.

  • Flexible 4–10 second pacingPick 4, 6, 8, or 10 seconds (default 6) to match hooks, mid-length demos, and fuller short narratives.

  • Landscape and portrait framingUse 16:9 for horizontal storytelling or 9:16 for Reels, Shorts, and TikTok-ready vertical layouts.

Parameters

ParameterRequirementDescription
promptRequired

String. Scene, action, camera, lighting, mood, and audio cues; 1–20,000 characters after trimming.

durationOptional

Integer. Output length in seconds; the Playground preselects 6.

Default64810
resolutionOptional

String. Output clarity tier; the Playground preselects 720p.

Default720p1080p4k
aspect_ratioOptional

String. Output framing; the Playground preselects 16:9.

Default16:99:16

How to Use

  1. Write the scene promptDescribe subject, setting, and main action in prompt—for example a product hero shot under golden-hour light with a confident walk-through.

  2. Add camera and audio directionSpecify push-in, dolly, or tracking moves, lighting mood, and dialogue or ambience cues so picture and sound stay aligned.

  3. Set output durationChoose 4, 6, 8, or 10 seconds (default 6) to match the pacing of the beat you want to generate.

  4. Select resolutionPick 720p for fast iteration, 1080p for clearer delivery, or 4K for high-detail finals.

  5. Choose aspect ratioSelect 16:9 landscape or 9:16 portrait to match the distribution layout.

  6. Review the cost and runCheck the cost shown on the Run button, finish the prompt and settings, then click Run.

  7. Preview and download the videoWhen the task finishes, preview picture and synced audio in the output panel, then select Download video to save the result.

Pricing

Billed per generation by duration and resolution tier, with native audio included. 1 credit = $0.005.

UsageRateDetails
720p / 1080p4s=120, 6s=150, 8s=200, 10s=220 creditsDefault 720p / 6s costs 150 credits ($0.75).
4k4s=250, 6s=300, 8s=350, 10s=450 credits4k / 6s costs 300 credits ($1.50).

Best Use Cases

  • Social short-form productionBatch 9:16 hooks and talking scenes with lip-synced audio for TikTok, Reels, and Shorts.

  • E-commerce product motionDescribe product reveals, texture close-ups, and ambient sound for shoppable demos and landing-page heroes.

  • Brand concept storyboardsTurn campaign briefs into 4–10 second cinematic beats to align creative direction before full production.

  • Dialogue and virtual-host clipsDirect speaking performances with language and tone cues so mouth motion tracks the scripted delivery.

Pro Tips

  • Structure the prompt as subject and setting, action, camera, lighting, then audio so each layer has a clear job.
  • Name camera moves explicitly—slow push-in, side tracking, or dolly—then separate subject motion into its own sentence.
  • End with an Audio line for dialogue language, music mood, and ambience tied to on-screen events.
  • State light quality and color—golden hour side light, soft window fill, neon reflections—to lock atmosphere.
  • Keep one primary action per clip and describe opening state, progression, and closing frame to reduce visual drift.

Usage notes

  • Gemini Omni Flash Text to Video is driven by a required prompt, with optional duration, resolution, and aspect_ratio.
  • Describe speech, music, or ambience in the prompt; native audio with lip sync is included in the result.
  • Prompt length is 1–20,000 characters after trimming; keep instructions concrete and cinematic.
  • After an API submission, save the returned task_id to query progress and retrieve the final media URL.

Related Models

Gemini Omni Flash Text to Video API frequently asked questions

What is the Gemini Omni Flash Text to Video API?

Gemini Omni Flash Text to Video is a Google DeepMind multimodal model for generating video from text prompts. It creates 4–10 second clips with native lip-synced audio, cinematic camera and lighting control, and resolution options from 720p to 4K. Built on Gemini’s unified multimodal architecture, it follows prompt physics and narrative pacing while keeping subject motion, scene continuity, and audiovisual alignment coherent. You can call it programmatically or try it from the playground above.

How long can Gemini Omni Flash Text to Video generate?

Each run supports 4, 6, 8, or 10 seconds, with the Playground preselecting 6 seconds. Choose shorter lengths for hooks and longer ones for fuller beats; exact credit cost scales with duration and resolution in the Pricing section.

Does Gemini Omni Flash Text to Video include native lip sync?

Yes. Dialogue, music, and ambience are generated with the picture and embedded in the MP4. Add language, tone, and environmental sound cues in the prompt so mouth motion and timing track the spoken delivery.

When should Gemini Omni Flash Text to Video use 4K?

Use 4K when finals need sharper texture, lighting reflections, and delivery clarity. For daily iteration, start at 720p or 1080p—same duration options, lower credit cost—then promote selected takes to 4K.

How does Gemini Omni Flash Text to Video follow camera prompts?

Write explicit moves such as slow push-in, dolly, or side tracking in their own sentences, separate from subject action. Pair them with lighting cues so framing and atmosphere stay intentional across the clip.

Is Gemini Omni Flash Text to Video suitable for 9:16 shorts?

Yes. Select 9:16 and describe vertical composition and subject placement for mobile feeds. Use 16:9 when you need landscape storytelling or horizontal ad placements.

How are Gemini Omni Flash Text to Video credits calculated?

Credits are charged per generation by duration and resolution. At 720p/1080p, 6 seconds costs 150 credits ($0.75); at 4K the same length costs 300 credits ($1.50). Full rate tables are in the Pricing section on this page.