One continuous ten-second fixed full-body shot inside a small circular circus rehearsal ring. A single adult male acrobat with closely cropped dark hair wears plain burgundy practice clothes and soft black shoes. He starts standing centered, holding one straight white baton horizontally. He tosses that single baton gently a short distance above his head, watches it, catches it cleanly with one hand, then makes a controlled half turn and gives a modest bow toward the camera. Keep his whole body and the baton visible throughout, with one consistent person and one consistent baton. Warm overhead rehearsal lighting, empty wooden benches, realistic movement, no cuts, no extra performers, no text or logos.
Grok Imagine Video Text to Video API
xai/grok-imagine-video/text-to-videoGrok Imagine Video Text to Video turns scene descriptions into short clips with text-directed action and camera movement. Prompts shape the scene’s subjects and motion, while selectable duration, visual mode, and aspect ratio set the format for each shot.
Get API KeyContinue with
Examples
REST API
Quick Start
Submit a request, save task_id, and retrieve the generated file when the task finishes.
Connect to Vidgo API
Store your API key on the server and send Authorization: Bearer VIDGO_API_KEY.
- Endpoint
- POST
https://api.vidgo.ai/api/generate/submit - Authentication
- Authorization: Bearer VIDGO_API_KEY
Submit a generation task
Set model and callback_url at the request root. Place the following fields inside input.
curl --request POST \
--url "https://api.vidgo.ai/api/generate/submit" \
--header "Authorization: Bearer $VIDGO_API_KEY" \
--header "Content-Type: application/json" \
--data '{
"model": "xai/grok-imagine-video/text-to-video",
"input": {
"prompt": "A ceramic cup on a wooden table. Steam rises gently as the camera slowly moves closer.",
"duration": 6,
"mode": "normal",
"aspect_ratio": "16:9"
}
}'Retrieve the result
Poll every 2–5 seconds while status is not_started or running. Stop at finished or failed.
Track status
GET https://api.vidgo.ai/api/generate/status/{task_id}Poll every 2–5 seconds while status is not_started or running. Stop at finished or failed.
not_startedrunningfinishedfailed{
"code": 200,
"data": {
"task_id": "example-task-id",
"status": "not_started",
"created_time": "2026-09-27T10:00:00Z"
}
}{
"code": 200,
"data": {
"task_id": "example-task-id",
"status": "finished",
"created_time": "2026-09-27T10:00:00Z",
"files": [
{
"file_url": "https://example.com/result.mp4",
"file_type": "video"
}
]
}
}Complete polling example
Submit a request, save task_id, and retrieve the generated file when the task finishes.
set -euo pipefail
: "${VIDGO_API_KEY:?Set VIDGO_API_KEY in your environment}"
REQUEST_BODY=$(cat <<'JSON'
{
"model": "xai/grok-imagine-video/text-to-video",
"input": {
"prompt": "A ceramic cup on a wooden table. Steam rises gently as the camera slowly moves closer.",
"duration": 6,
"mode": "normal",
"aspect_ratio": "16:9"
}
}
JSON
)
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
--request POST \
--url "https://api.vidgo.ai/api/generate/submit" \
--header "Authorization: Bearer $VIDGO_API_KEY" \
--header "Content-Type: application/json" \
--data "$REQUEST_BODY")
TASK_ID=$(printf '%s' "$SUBMIT_RESPONSE" | jq -r '.data.task_id // .task_id // empty')
if [ -z "$TASK_ID" ]; then
printf 'Submit response did not include task_id:
%s
' "$SUBMIT_RESPONSE" >&2
exit 1
fi
while true; do
STATUS_RESPONSE=$(curl --silent --show-error --fail-with-body \
--url "https://api.vidgo.ai/api/generate/status/$TASK_ID" \
--header "Authorization: Bearer $VIDGO_API_KEY")
STATUS=$(printf '%s' "$STATUS_RESPONSE" | jq -r '.data.status // .status // empty')
case "$STATUS" in
finished)
printf '%s' "$STATUS_RESPONSE" | jq -r '(.data.files // .files // [])[]?.file_url'
break
;;
failed)
printf '%s' "$STATUS_RESPONSE" | jq -r '.data.error_message // .error_message // "Generation failed"' >&2
exit 1
;;
not_started|running)
sleep 2
;;
*)
printf 'Unexpected task status: %s
' "$STATUS" >&2
exit 1
;;
esac
doneRequest Parameters
Set model and callback_url at the request root. Place the following fields inside input.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
| input.prompt | string | Yes | – | Describe the desired scene and motion. Enter 1–5,000 characters. |
| input.aspect_ratio | string | No | – | Choose 1:1, 2:3, 3:2, 16:9, or 9:16. |
| input.mode | string | No | – | Generation style: fun, normal, or spicy. |
| input.duration | integer | No | 6 | Video duration in seconds. Choose 6 or 10; defaults to 6. |
Response Fields
The submit response returns task_id. The status response returns generated files or the task error.
| Field | Type | Description |
|---|---|---|
| code | integer | Response code; 0 or 200 indicates success. |
| data.task_id | string | Task identifier for status queries. |
| data.status | string | not_started, running, finished, failed |
| data.created_time | string | Task creation timestamp. |
| data.files[].file_url | string | URL of the generated file. |
| data.files[].file_type | string | video |
| data.error_message | string | null | Reason for a failed task. |
Task Lifecycle
Poll every 2–5 seconds while status is not_started or running. Stop at finished or failed.
not_startedAccepted and queued.
runningGeneration in progress.
finishedRead the generated files.
failedRead error_message and stop polling.
Polling and Errors
- PollingPoll every 2–5 seconds while status is not_started or running. Stop at finished or failed.
- Retry a status queryBack off on 429 and temporary server errors, then query the same task_id.
- CallbackProvide a public callback_url at the request root to receive completion notifications.
Model Specifications
| Specification | Value | Details |
|---|---|---|
| Input | Prompt | Describe the desired scene and motion. Enter 1–5,000 characters. |
| Output | video | 6- or 10-second video. |
| aspect_ratio | 1:1 · 2:3 · 3:2 · 16:9 · 9:16 | Choose 1:1, 2:3, 3:2, 16:9, or 9:16. |
| mode | fun · normal · spicy | Generation style: fun, normal, or spicy. |
| duration | 6 · 10 | Video duration in seconds. Choose 6 or 10; defaults to 6. |
Related Models
Grok Imagine Video Text to Video API frequently asked questions
What is the Grok Imagine Video Text to Video API?
Grok Imagine Video Text to Video is an xAI model for creating video scenes from text. It generates 6- or 10-second clips with selectable visual modes and square, portrait, or landscape framing. The text describes the subject, setting, action, and camera direction together, giving the shot a defined scene and movement. You can call it through the API or try it in the Playground tab.
How does Grok Imagine Video Text to Video use camera instructions?
Include the intended camera movement alongside the subject’s action in prompt. For example, describe a camera moving closer to a stationary object or following a person through a scene, so subject motion and camera motion have separate roles.
Which visual modes does Grok Imagine Video Text to Video offer?
The mode field accepts fun, normal, and spicy. Select a mode alongside a scene prompt; when comparing treatments, keep the prompt and other settings unchanged to evaluate that choice.
How do I choose the duration in Grok Imagine Video Text to Video?
Set duration to 6 or 10 seconds; omitting it selects 6 seconds. Plan the action around the chosen length, using a focused movement for a short shot or more time for an action to unfold.
Which frames can Grok Imagine Video Text to Video generate?
Set aspect_ratio to 1:1, 2:3, 3:2, 16:9, or 9:16. This sets the output frame for a square post, portrait scene, or landscape composition.
When should I choose Grok Imagine Video Text to Video instead of Grok Imagine Video Image to Video?
Choose Grok Imagine Video Text to Video when the scene, subjects, and action are described in a written brief. Choose Grok Imagine Video Image to Video when an existing picture should provide the visual starting point.
What does a 10-second Grok Imagine Video Text to Video clip cost?
A 10-second clip costs 40 credits ($0.200), and a 6-second clip costs 30 credits ($0.150). Billing is per generated video at the selected duration.















