Grok Imagine Video Text to Video API

xai/grok-imagine-video/text-to-video
6 / 10 seconds · 3 modes

Grok Imagine Video Text to Video turns scene descriptions into short clips with text-directed action and camera movement. Prompts shape the scene’s subjects and motion, while selectable duration, visual mode, and aspect ratio set the format for each shot.

Get API Key
Input
664/5000

Examples

One continuous ten-second fixed full-body shot inside a small circular circus rehearsal ring. A single adult male acrobat with closely cropped dark hair wears plain burgundy practice clothes and soft black shoes. He starts standing centered, holding one straight white baton horizontally. He tosses that single baton gently a short distance above his head, watches it, catches it cleanly with one hand, then makes a controlled half turn and gives a modest bow toward the camera. Keep his whole body and the baton visible throughout, with one consistent person and one consistent baton. Warm overhead rehearsal lighting, empty wooden benches, realistic movement, no cuts, no extra performers, no text or logos.

A single continuous six-second aerial tracking shot from behind and slightly above one small cream-yellow minibus driving forward at a steady sensible speed along a winding paved road through lush rolling green hills. Follow the same vehicle smoothly as it rounds one broad S-shaped bend. Keep its wheels on the road, preserve the road geometry, vehicle shape and forward direction. Early morning sunlight, gentle long shadows, distant layered hills, natural documentary color. No abrupt zoom, no cut, no other vehicles, no text, branding or advertising.

Grok Imagine Video Text to Video

Grok Imagine Video Text to Video is xAI’s model for creating short video scenes from a written description. The prompt combines a subject, setting, action, and camera direction into a moving scene, while generation modes offer another way to explore its visual treatment. Selectable clip lengths and aspect ratios suit storyboards, product reveal concepts, and vertical social scenes developed directly from a creative brief.

Why Choose Grok Imagine Video Text to Video?

  • From a written scene to a moving shotA scene brief can describe both what appears and how it moves, making this endpoint useful for exploring a shot before a source image has been chosen.

  • Subject action and camera direction togetherThe prompt can describe an object’s movement alongside a pan, push-in, or other camera direction, connecting the subject’s action to the intended framing.

  • Alternative visual treatmentsThe fun, normal, and spicy modes provide selectable generation styles alongside the scene description, so the same idea can be explored with different settings.

  • Short clips for different placementsTwo clip lengths and five aspect ratios fit individual storyboard shots, landscape scenes, and vertical social concepts.

Parameters

ParameterRequirementDescription
promptRequired

Describe the desired scene and motion. Enter 1–5,000 characters.

aspect_ratioOptional

Choose 1:1, 2:3, 3:2, 16:9, or 9:16.

1:12:33:216:99:16
modeOptional

Generation style: fun, normal, or spicy.

funnormalspicy
durationOptional

Video duration in seconds. Choose 6 or 10; defaults to 6.

Default610

How to Use

  1. Describe the scene and movementEnter the subject, setting, action, and camera direction.

  2. Set the clip formatChoose Duration, Mode, and Aspect Ratio.

  3. Run and view the resultReview the price and select Run, then preview and download the completed result.

Pricing

One credit equals $0.005.

UsageRateDetails
6-second video30 credits / video$0.150 / video
10-second video40 credits / video$0.200 / video

Use Cases

  • Storyboard motion studiesTurn a written shot description into a short moving scene to discuss subject action and camera direction before production.

  • Product reveal conceptsDescribe an imagined product, its surroundings, and a reveal movement to explore the timing and framing of a short introduction.

  • Vertical story scenesCreate a self-contained action in a portrait frame for a short social narrative, starting from the script’s scene description.

Prompt Tips

  • Describe the subject and setting first, then give the subject a clear action.
  • Distinguish subject motion from camera motion, such as “the cyclist rides past as the camera pans right.”
  • Plan one focused shot around the selected 6- or 10-second duration.
  • Choose aspect_ratio for the intended placement and describe where the action sits within that frame.
  • When comparing modes, keep the scene prompt, duration, and aspect_ratio the same so the setting change is easy to assess.

Usage Notes

  • Save the task_id to retrieve the result, or provide callback_url for a completion notification.

Related Models

Grok Imagine Video Text to Video API frequently asked questions

What is the Grok Imagine Video Text to Video API?

Grok Imagine Video Text to Video is an xAI model for creating video scenes from text. It generates 6- or 10-second clips with selectable visual modes and square, portrait, or landscape framing. The text describes the subject, setting, action, and camera direction together, giving the shot a defined scene and movement. You can call it through the API or try it in the Playground tab.

How does Grok Imagine Video Text to Video use camera instructions?

Include the intended camera movement alongside the subject’s action in prompt. For example, describe a camera moving closer to a stationary object or following a person through a scene, so subject motion and camera motion have separate roles.

Which visual modes does Grok Imagine Video Text to Video offer?

The mode field accepts fun, normal, and spicy. Select a mode alongside a scene prompt; when comparing treatments, keep the prompt and other settings unchanged to evaluate that choice.

How do I choose the duration in Grok Imagine Video Text to Video?

Set duration to 6 or 10 seconds; omitting it selects 6 seconds. Plan the action around the chosen length, using a focused movement for a short shot or more time for an action to unfold.

Which frames can Grok Imagine Video Text to Video generate?

Set aspect_ratio to 1:1, 2:3, 3:2, 16:9, or 9:16. This sets the output frame for a square post, portrait scene, or landscape composition.

When should I choose Grok Imagine Video Text to Video instead of Grok Imagine Video Image to Video?

Choose Grok Imagine Video Text to Video when the scene, subjects, and action are described in a written brief. Choose Grok Imagine Video Image to Video when an existing picture should provide the visual starting point.

What does a 10-second Grok Imagine Video Text to Video clip cost?

A 10-second clip costs 40 credits ($0.200), and a 6-second clip costs 30 credits ($0.150). Billing is per generated video at the selected duration.