Gemini Omni 1.1 Flash Text-to-Video API

google/gemini-omni-1.1-flash/text-to-video

Gemini Omni 1.1 Flash (Text-to-Video) turns text prompts into short videos with native audio, camera and lighting control, and output from 360p to 4K. Direct scenes, action, and sound in one creative brief to produce 4–10 second clips for product concepts, storyboards, and social content.

Input
564/20000
Output
Idle

Your generated video will appear here

Add your prompt and required media, review the settings, then click Run.

720p · 8 sec = $0.375 (75 credits)
Continue with

Examples

One continuous 8-second cinematic shot. Slow push-in down a wet-asphalt night avenue lined with neon storefronts; rain beads on the pavement and a metal awning. No people in close-up, no shops selling products, no readable signs. Lighting: high-contrast neon magenta and teal bouncing off black puddles. Camera: locked-axis slow dolly-in at chest height, no cuts. Native audio: rain on metal, distant traffic hiss, and a low analog music pulse. No logos, no readable text, no products, no packaging, no prices, no CTA, no advertising, no watermark, no brand marks.

Vertical 9:16, 4 seconds. Dawn on a weathered wooden fishing pier, pale peach sky, no rain, no city. One adult fisherman in a plain grey jacket stands at the rail, backs three-quarter to camera, and says one short quiet line in English: "The tide is turning." Camera: gentle handheld hold, slight drift, no cuts. Lighting: cool marine dawn with a warm horizon band. Native audio: one spoken line, sea wind, water slapping pilings. Not a cafe, shop, or product scene. No logos, no readable text, no products, no packaging, no prices, no CTA, no advertising, no watermark, no brand marks.

One continuous 4-second 16:9 macro shot. A single clear glass marble rolls along a brushed-steel toy track, tapping each joint. Low camera two centimeters above the rail, tracking beside the marble. Lighting: hard workshop sidelight, sharp highlights on glass and metal. Native audio: precise Foley of glass on steel, tiny ticks, quiet room tone. No hands, no toys as products, no packaging. No logos, no readable text, no products, no packaging, no prices, no CTA, no advertising, no watermark, no brand marks.

Gemini Omni 1.1 Flash Text-to-Video

Gemini Omni 1.1 Flash Text-to-Video turns text prompts into short videos with speech, music, and ambient sound. Describe subject action, camera movement, lighting, and audio to bring product ideas, story scenes, and brand atmospheres into moving footage.

Why Choose This?

  • Text-driven scene creationBuild a scene through descriptions of subjects, surroundings, and action, turning a product idea or story moment into a watchable clip.

  • Camera movement controlDescribe push-ins, tracking shots, orbits, and shot sizes to direct attention toward product details or key story moments.

  • Lighting and atmosphereSpecify light sources, color, and mood to explore morning light, neon streets, or soft interiors around the same creative idea.

  • Native speech and soundCreate dialogue, music, and ambient sound alongside the video, planning visual content and its soundtrack in one prompt.

  • Multiple output resolutionsChoose 360p, 720p, 1080p, or 4K to match the output specifications for previews, editing, and content delivery.

  • Landscape and portrait clipsCombine 16:9 or 9:16 framing with a 4, 6, 8, or 10 second duration to plan composition and pacing for landscape stories or portrait social content.

Parameters

ParameterRequirementDescription
promptRequired

String. Defines the scene, action, camera, lighting, mood, and sound; 1–20,000 characters after trimming.

durationOptional

Integer. Sets output length; the Playground preselects 8 seconds.

Default84610
resolutionOptional

String. Sets output resolution; the Playground preselects 720p.

Default720p360p1080p4k
aspect_ratioOptional

String. Controls output framing; the Playground preselects 16:9.

Default16:99:16

How to Use

  1. Describe the scene and actionEnter a subject, setting, and main action, such as a glass marble rolling along a metal track beside a warm desk lamp.

  2. Add camera and sound directionDescribe camera movement, lighting, and speech, music, or ambient sound, such as a low tracking shot accompanied by crisp rolling sounds.

  3. Choose a resolutionChoose 360p, 720p, 1080p, or 4K, with 720p selected by default, to match the output specifications for editing or presentation.

  4. Choose a durationChoose 4, 6, 8, or 10 seconds, with 8 seconds selected by default, to plan the clip around its main action and pacing.

  5. Choose an aspect ratioChoose 16:9 landscape or 9:16 portrait, with 16:9 selected by default, and frame the subject for the intended layout.

  6. Review the cost and runReview the cost shown on the Run button, complete the required uploads and prompt, then click Run.

  7. Preview and download the videoWhen the task finishes, preview the video and sound in the output panel, then select Download video to save the result.

Pricing

Billed per generation based on video duration and resolution tier, with native audio included in the result. 1 credit = $0.005.

UsageRateDetails
360p / 720p / 1080p4s=45, 6s=60, 8s=75, 10s=90 creditsDefault 720p / 8s costs 75 credits ($0.375).
4k4s=105, 6s=120, 8s=135, 10s=150 credits4k / 8s costs 135 credits ($0.675).

Best Use Cases

  • Product concept filmsDescribe a product scene, key action, and musical direction to create concept footage for creative pitches and brand presentations.

  • Storyboard previewsTurn one action beat from a script into a clip with sound to communicate shot size, camera movement, and staging.

  • Portrait brand contentBuild a scene around one brand theme and choose 9:16 framing to create portrait footage for social channels.

  • Atmospheric footageCombine rain, streets, or interior lighting with ambient sound to create clips for an opening sequence or story transition.

Pro Tips

  • Organize the prompt around subject and setting, action, camera, lighting, and audio so each part has a clear creative role.
  • Describe subject action and camera movement separately, such as “The person walks toward the door. The camera tracks slowly from the side,” to define how they relate.
  • For a continuous shot, specify the starting camera position, direction of travel, and final composition.
  • Introduce speech, music, and effects with “Audio:” and connect sounds to actions, such as a hinge creaking as a door opens.
  • Build each 4–10 second clip around one main action, describing its opening state, progression, and closing image.

Usage notes

  • Gemini Omni 1.1 Flash Text-to-Video generates video from a required text prompt (prompt), with duration, resolution, and aspect ratio settings to configure the output.
  • Describe speech, music, or ambient sound in the prompt; native audio is included in the result.
  • Provide a prompt of 1–20,000 characters after trimming, describing the scene, action, camera, and sound.
  • Save the task_id returned by an API submission to query progress and retrieve the result.

Gemini Omni 1.1 Flash Text-to-Video API frequently asked questions

What is the Gemini Omni 1.1 Flash Text-to-Video API?

Gemini Omni 1.1 Flash Text-to-Video is a Google model for generating video from text. It creates short videos up to 4K resolution with native dialogue, music, and ambient sound effects from text prompts, supporting precise camera movement, lighting, and atmospheric scene control. Built on Gemini's multimodal intelligence architecture, it strictly follows prompt physics and visual continuity while delivering cinematic narrative pacing. You can call it programmatically or try it from the playground above.

Can Gemini Omni 1.1 Flash Text-to-Video generate speech and music?

Yes. The model features native audio-visual synchronization without requiring post-production audio synthesis. By describing dialogue lines, musical moods, or ambient sound effects (such as footsteps or raindrops) in your prompt, the model embeds synchronized audio tracks directly into the output MP4 video.

What is the maximum duration for a single Gemini Omni 1.1 Flash Text-to-Video generation?

A single call supports generating clips up to 10 seconds, with choices of 4, 6, 8, or 10 seconds (the playground defaults to 8 seconds). For narratives requiring broader story arcs, you can structure sequential story beats cleanly within the prompt or extend shots progressively using continuation workflows.

Does Gemini Omni 1.1 Flash Text-to-Video support 4K resolution?

Yes. The model provides output resolution tiers from 360p, 720p, and 1080p up to 4K. 4K mode is ideal for high-end commercial showcases and final delivery by preserving fine lighting reflections and surface textures, whereas 720p offers a balanced preview experience for daily iteration.

Can I use 360p to rapidly prototype Gemini Omni 1.1 Flash Text-to-Video concepts?

Yes. The 360p resolution is designed specifically for low-latency drafting, significantly accelerating generation speed and reducing credit consumption. You can validate staging, subject movement, and camera timing in 360p drafts before switching to 1080p or 4K for production rendering.

How do I direct camera movement in Gemini Omni 1.1 Flash Text-to-Video?

We recommend separating subject choreography and camera directions into distinct sentences within your prompt. Using explicit cinematographic terms such as slow dolly forward, 360-degree orbit, or low-angle tracking shot, combined with specific lighting cues (like warm side lighting or neon reflections), produces controlled and coherent motion.

Is Gemini Omni 1.1 Flash Text-to-Video suitable for 9:16 vertical videos?

Yes. The model supports standard 16:9 landscape and 9:16 portrait aspect ratios. When selecting 9:16, focus prompt descriptions on vertical composition and subject positioning, allowing the model to naturally optimize focal balance and background framing for mobile feeds.