Grok Imagine Image 2.0 Text to Image API

xai/grok-imagine-image-v2.0/text-to-image
Low/Medium · 1K/2K · 5 ratios

Grok Imagine Image 2.0 Text to Image transforms text prompts into high-fidelity images up to 2K resolution, featuring designer-grade typography, meticulous detail control, and 5 standard aspect ratios. It adheres faithfully to complex multi-layered prompts and spatial arrangements while maintaining crisp legibility and stylistic consistency across dense layouts.

Get API Key
Input
1007/8000
OutputReady
1 × 12 = 12 credits · $0.060

Continue with

Examples

output.jpg

Create a cinematic but physically believable underwater exploration photograph inside a vast limestone cenote. A single scuba diver with a black wetsuit, one silver tank, mask and yellow fins swims horizontally from the lower left toward a sunlit opening in the upper right, small enough to reveal the scale of the cavern. Show exactly one diver with two arms and two fins, natural buoyancy, a subtle trail of rising air bubbles. Several sharp shafts of pale turquoise sunlight enter through the broken rock ceiling, revealing tiny suspended particles; faint caustic patterns dance over the nearest limestone ledge. The foreground rock is richly textured and dark, the middle-distance diver is crisp, and the distant water fades into deep cobalt blue. Wide 16:9 environmental composition, nuanced highlights rather than neon glow, credible underwater optics and equipment, a sense of discovery and quiet. A full-frame photograph without titles, captions, signs, branding, advertising, borders or watermarks.

output.jpg

A candid 35mm documentary photograph of a quiet outdoor Go game in a lived-in rural courtyard. Exactly three elderly adults: one woman seated at the left of a worn wooden table, one man seated at the right facing her, and one neighbor standing a step behind the table watching thoughtfully. The two players rest their hands naturally near the edge of the table; no one is holding anything. A single Go board with black and white stones lies flat between them, with two small stone bowls beside it. A broad old tree casts irregular dappled afternoon light across their weathered faces, cotton shirts, stools and dusty stone paving. In the background, an ordinary plaster house, a half-open wooden door and climbing vines are softly out of focus. Warm, unposed human interaction, realistic age and skin texture, believable hand anatomy, clear spatial relationships, subtle film grain, landscape 3:2 composition. No spectators beyond these three people, no staged commercial aesthetic, no writing, logos, slogans, advertising, borders or watermark.

output.jpg

A richly textured hand-painted fantasy story illustration in a vertical 2:3 frame. One tiny human traveler in an indigo hooded cloak shelters beneath a single enormous ochre mushroom during a rainy forest night. The traveler sits on a mossy root in the lower center, holding one small warm amber lantern; show the complete figure. The mushroom cap fills the upper third like a protective roof, with delicate radial gills visible underneath and heavy droplets hanging from its rim. The lantern gently lights the traveler's hands, a small canvas satchel, damp moss and nearby roots, while the deeper forest recedes into layered blue-green silhouettes. Fine silver rain falls outside the shelter, with tiny ripples in shallow puddles. Maintain a coherent miniature scale: the mushroom is much taller than the traveler, the moss looks like a small landscape. Painterly gouache and colored-pencil edges, tactile brushwork, restrained luminous contrast, tender narrative atmosphere. No animals, no second traveler, no text, signs, promotional elements, borders or watermark.

Grok Imagine Image 2.0 Text to Image Overview

Grok Imagine Image 2.0 Text to Image is a state-of-the-art high-fidelity text-to-image generation model engineered by xAI. Ranking among the global leaders on the Text-to-Image Arena benchmark, it excels at commercial graphic design, razor-sharp typography rendering, granular prompt adherence, and tactile physical realism across 1K and 2K resolutions and 5 aspect ratios.

Why Choose Grok Imagine Image 2.0 Text to Image?

  • Designer-Grade TypographyOvercomes legacy text generation limitations to produce crisp, legible headers, small print, and intricate poster lettering directly in the image.

  • Granular Prompt AdherenceFaithfully maps complex instructions concerning subject relationships, lighting directions, textures, and spatial layouts without omitting key details.

  • Pristine 2K High ResolutionDelivers pristine outputs across 1K and 2K tiers with configurable Low and Medium quality settings, balancing processing velocity with fine texture fidelity.

  • Versatile Native Aspect RatiosGenerates natively in 1:1, 2:3, 3:2, 9:16, and 16:9 framing, tailored directly for social campaigns, commercial posters, and cinematic previsualization.

  • Predictable Billing & Automatic RefundsCharges transparently per output image according to tier rates, with automatic refunds if a generation task encounters an unexpected interruption.

Parameters

ParameterRequirementDescription
promptRequired

Prompt must be a non-blank string of 1–8,000 characters.

nOptional

Integer from 1 to 4. Default: 1.

Default1234
aspect_ratioOptional

Supports 1:1, 2:3, 3:2, 9:16, 16:9. Default: 1:1.

Default1:12:33:29:1616:9
resolutionOptional

Choose 1K or 2K. Default: 1K.

Default1K2K
qualityOptional

Low or Medium. Default: Medium.

Defaultmediumlow

How to Use

  1. Draft a structured promptDescribe core subjects, spatial staging, lighting mood, and specific typography elements (surround desired lettering with double quotes, up to 8,000 characters).

  2. Select resolution and qualityPick 1K or 2K resolution, and configure Low or Medium quality to match your visual fidelity and performance requirements.

  3. Configure ratio and countChoose the aspect ratio suited to your target channel (1:1, 2:3, 3:2, 9:16, or 16:9) and set the number of outputs (n from 1 to 4).

  4. Submit generation jobDispatch your API request to receive a distinct task ID, beginning asynchronous execution across the distributed compute infrastructure.

  5. Retrieve verified assetsQuery the task status endpoint until finished, then access the signed download URLs for your verified PNG assets.

Pricing

Total credits = output count × tier rate. 1 credit = $0.005.

UsageRateDetails
Low / 1K8 credits / output image$0.040 / image
Medium / 1K12 credits / output image$0.060 / image
Low / 2K12 credits / output image$0.060 / image
Medium / 2K16 credits / output image$0.080 / image

Best Use Cases

  • Brand marketing posters and visual adsLeverage advanced typographic clarity to produce commercial collateral featuring readable slogans and marketing layouts.

  • UI mockups and product concept designsGenerate application screens, graphic kits, and user interface concepts with crisp iconography and structural coherence.

  • E-commerce product visual assetsOutline physical materials and studio illumination setups to create studio-grade product hero shots and contextual imagery.

  • Cinematic concept art and world-buildingTurn intricate narrative scripts into sweeping, atmosphere-rich concept paintings with convincing physical illumination.

Pro Tips

  • Structure prompts hierarchically: subject appearance, composition framing, lighting atmosphere, material texture, and graphic text elements.
  • When generating specific text inside the visual, wrap the exact target wording in straight double quotes, such as "VISIT MARS".
  • Specify explicit lighting setups to elevate realism, such as soft studio bounce illumination, dramatic Rembrandt side lighting, or warm golden-hour rim highlights.
  • Leverage the n parameter to produce up to 4 distinct variants simultaneously, enabling swift creative comparison under identical prompt parameters.

Notes

  • Prompt is required and must contain 1 to 8,000 characters after removing leading and trailing whitespace.
  • Output count n must be an integer between 1 and 4, inclusive.
  • Generation is asynchronous; track progress via task status polling or configure a callback_url for automatic payload delivery upon completion.

Related Models

Grok Imagine Image 2.0 Text to Image API frequently asked questions

What is the Grok Imagine Image 2.0 Text to Image API?

Grok Imagine Image 2.0 Text to Image is an advanced text-to-image synthesis model developed by xAI. It transforms descriptive text prompts into high-fidelity visuals up to 2K resolution, featuring designer-grade typography, nuanced prompt adherence, and support for 5 native aspect ratios. Engineered with unified multimodal architecture, it faithfully preserves intricate spatial and textural directions while keeping small print legible and visually stable. You can call it programmatically or try it from the playground above.

Does Grok Imagine Image 2.0 Text to Image support readable in-image text?

Yes. The model is specifically optimized for typography and complex multi-part visual layouts, producing crisp, legible headings, small print, and signage. Enclosing the target phrase in straight double quotes within your prompt yields the most accurate typographic rendering.

What resolutions and aspect ratios does Grok Imagine Image 2.0 Text to Image support?

It supports 1K and 2K output resolutions across 1:1, 2:3, 3:2, 9:16, and 16:9 aspect ratios. Visual compositions and illumination schemes are generated directly in the requested framing without post-generation cropping.

Can I generate multiple images in a single API call?

Yes. By setting the n parameter between 1 and 4, the model outputs up to 4 distinct compositional variants in a single job, simplifying creative review and selection.

How is Grok Imagine Image 2.0 Text to Image billed?

Pricing is calculated per output image based on resolution and quality tier: 8 credits ($0.040) for 1K Low, 12 credits ($0.060) for 1K Medium or 2K Low, and 16 credits ($0.080) for 2K Medium. Tasks that encounter an unexpected failure receive an immediate automatic full refund.

How does Grok Imagine Image 2.0 Text to Image handle long prompts?

With a generous prompt capacity of up to 8,000 characters, the model parses multi-clause instructions concerning subject interactions, material reflections, background depth, and specific stylistic cues with high fidelity.

What is the best way to structure prompts for Grok Imagine Image 2.0 Text to Image?

Organize your prompt from core subject details to spatial placement, lighting, atmosphere, and visual style. Clear descriptive separation helps the generation engine establish balanced focal depth and realistic rendering.