Grok Imagine Image 2.0 Text to Image API

xai/grok-imagine-image-v2.0/text-to-image
Low/Medium · 1K/2K · 5 种比例

Grok Imagine Image 2.0 Text to Image 将纯文本提示词转化为最高 2K 分辨率的高保真图像,支持设计级排版文字渲染、微小细节精准把控,以及 5 种主流画面比例。它能够严格遵循复杂的多层级提示词描述与空间构图关系,并在渲染小字排版和密集视觉元素时保持清晰连贯与风格统一。

获取 API Key
输入
1007/8000
输出已就绪
1 × 12 = 12 积分 · $0.060

继续使用

示例

output.jpg

Create a cinematic but physically believable underwater exploration photograph inside a vast limestone cenote. A single scuba diver with a black wetsuit, one silver tank, mask and yellow fins swims horizontally from the lower left toward a sunlit opening in the upper right, small enough to reveal the scale of the cavern. Show exactly one diver with two arms and two fins, natural buoyancy, a subtle trail of rising air bubbles. Several sharp shafts of pale turquoise sunlight enter through the broken rock ceiling, revealing tiny suspended particles; faint caustic patterns dance over the nearest limestone ledge. The foreground rock is richly textured and dark, the middle-distance diver is crisp, and the distant water fades into deep cobalt blue. Wide 16:9 environmental composition, nuanced highlights rather than neon glow, credible underwater optics and equipment, a sense of discovery and quiet. A full-frame photograph without titles, captions, signs, branding, advertising, borders or watermarks.

output.jpg

A candid 35mm documentary photograph of a quiet outdoor Go game in a lived-in rural courtyard. Exactly three elderly adults: one woman seated at the left of a worn wooden table, one man seated at the right facing her, and one neighbor standing a step behind the table watching thoughtfully. The two players rest their hands naturally near the edge of the table; no one is holding anything. A single Go board with black and white stones lies flat between them, with two small stone bowls beside it. A broad old tree casts irregular dappled afternoon light across their weathered faces, cotton shirts, stools and dusty stone paving. In the background, an ordinary plaster house, a half-open wooden door and climbing vines are softly out of focus. Warm, unposed human interaction, realistic age and skin texture, believable hand anatomy, clear spatial relationships, subtle film grain, landscape 3:2 composition. No spectators beyond these three people, no staged commercial aesthetic, no writing, logos, slogans, advertising, borders or watermark.

output.jpg

A richly textured hand-painted fantasy story illustration in a vertical 2:3 frame. One tiny human traveler in an indigo hooded cloak shelters beneath a single enormous ochre mushroom during a rainy forest night. The traveler sits on a mossy root in the lower center, holding one small warm amber lantern; show the complete figure. The mushroom cap fills the upper third like a protective roof, with delicate radial gills visible underneath and heavy droplets hanging from its rim. The lantern gently lights the traveler's hands, a small canvas satchel, damp moss and nearby roots, while the deeper forest recedes into layered blue-green silhouettes. Fine silver rain falls outside the shelter, with tiny ripples in shallow puddles. Maintain a coherent miniature scale: the mushroom is much taller than the traveler, the moss looks like a small landscape. Painterly gouache and colored-pencil edges, tactile brushwork, restrained luminous contrast, tender narrative atmosphere. No animals, no second traveler, no text, signs, promotional elements, borders or watermark.

Grok Imagine Image 2.0 Text to Image 概览

Grok Imagine Image 2.0 Text to Image 是 xAI 研发的新一代高保真文本生成图像模型。它在权威评测基准 Text-to-Image Arena 中高居全球前列,以商业级排版设计、微小文本清晰呈现、深度提示词细节遵循与真实物理质感为核心特色,提供 1K 与 2K 分辨率档位及 1:1、2:3、3:2、9:16、16:9 五种画幅比例。

为什么选择 Grok Imagine Image 2.0 Text to Image?

  • 设计级排版与文字渲染突破传统模型难以生成清晰文字的痛点,原生支持海报标语、信息图文与小字排版的高清锐利输出。

  • 卓越的提示词细节遵循精确解析复杂指令中的主体属性、材质质感、环境光线与空间透视关系,避免主体混淆或关键要素遗漏。

  • 高保真 2K 高清输出支持 1K 与 2K 两种分辨率,配合 Low 与 Medium 画质档位,在渲染速度与细节密度之间提供理想平衡。

  • 多构图画幅自由适配原生支持 1:1、2:3、3:2、9:16 和 16:9 五种构图比例,直接适配社交媒体、商业海报与影视概念设计需求。

  • 透明计费与失败自动退还按输出张数与分辨率阶梯精确扣除点数,若任务因不可抗力中断,系统全额自动返还扣除积分。

参数

参数要求说明
prompt必填

提示词必须为 1–8,000 个字符的非空字符串。

n可选

整数 1–4,默认 1。

默认1234
aspect_ratio可选

支持 1:1、2:3、3:2、9:16、16:9,默认 1:1。

默认1:12:33:29:1616:9
resolution可选

选择 1K 或 2K,默认 1K。

默认1K2K
quality可选

Low 或 Medium,默认 Medium。

默认mediumlow

使用教程

  1. 编写结构化提示词详细描述画面主体、构图布局、艺术风格与特定排版文字(可用双引号包裹文字内容,最多 8,000 字符)。

  2. 选择分辨率与画质根据使用场景选择 1K 或 2K 分辨率,并指定 Low 或 Medium 画质档位以匹配精度要求。

  3. 指定画面比例与张数选择适合目标媒介的画幅比例(1:1、2:3、3:2、9:16 或 16:9),并设置生成张数(1–4 张)。

  4. 发起异步生成任务发送 API 请求提交任务并立即获取唯一任务 ID,由后台分布式算力集群异步执行渲染。

  5. 获取高清成品图像轮询任务状态接口至完成状态,读取生成的各张高保真 PNG 图像下载链接。

价格

总积分 = 输出数量 × 输出档位单价。1 积分 = $0.005。

用量费率详情
Low / 1K8 积分 / 输出图片$0.040 / 张
Medium / 1K12 积分 / 输出图片$0.060 / 张
Low / 2K12 积分 / 输出图片$0.060 / 张
Medium / 2K16 积分 / 输出图片$0.080 / 张

最佳应用场景

  • 品牌营销海报与图文广告利用顶尖的文字渲染与排版能力,直接生成包含清晰品牌标语与活动排版的宣发物料。

  • UI 概念原型与产品视觉生成带有清晰界面元素、应用图标与排版信息的应用界面视觉提案与产品概念图。

  • 电商商品摄影与场景渲染描述产品质感与特定布光环境,批量产出具有摄影级真实感的商用主图与氛围图。

  • 影视概念美术与世界观搭建将复杂的世界观与场景细节文本转化为极具视觉冲击力与叙事张力的高清概念艺术原画。

专家技巧

  • 建议遵循“核心主体 + 空间构图 + 光影氛围 + 风格材质 + 排版文字”的递进结构组织提示词,帮助模型准确定位视觉重心。
  • 画面中需要精准呈现的具体英文字符,建议使用英文字符半角双引号括起,例如 "VISIT MARS"。
  • 如果需要精细的光线质感,可在提示词中指定具体布光类型,如柔和摄影棚漫反射光、强烈伦勃朗侧光或黄昏逆光边缘轮廓光。
  • 单次请求可通过设置 n 参数生成多达 4 张不同构图与细节的变体图像,便于在同一提示词下快速遴选最佳方案。

注意事项

  • 提示词为必填参数,去除首尾空白后字符数必须在 1 至 8,000 字符之间。
  • 单次请求的输出图片张数 n 必须为 1 至 4 之间的整数。
  • 任务执行采用异步机制,提交请求后立即返回 task_id,建议通过轮询或指定 callback_url 回调接收最终结果。

相关模型推荐

Grok Imagine Image 2.0 Text to Image API 常见问题

Grok Imagine Image 2.0 Text to Image API 是什么?

Grok Imagine Image 2.0 Text to Image 是 xAI 研发的高保真文本生成图像模型。它根据文本提示词直接生成最高 2K 分辨率的高清图像,具备专业级排版文字渲染、微小细节精准呈现与 5 种主流构图画幅支持。基于前沿多模态深度架构,它在严格遵循多层级空间与材质提示词的同时,确保排版小字清晰可读且画质稳定细腻。你可以通过 API 进行程序化调用,也可以在上方体验区直接在线试用。

Grok Imagine Image 2.0 Text to Image 支持在画面中生成可读文字吗?

支持。该模型专门优化了排版与复杂视觉布局训练,能够在海报、招牌、包装与信息图中生成字迹清晰锐利的小字和标题文本。在提示词中使用英文双引号包裹目标文字能够获得最准确的呈现。

Grok Imagine Image 2.0 Text to Image 支持哪些输出分辨率与画幅比例?

支持 1K 与 2K 两种分辨率,以及 1:1、2:3、3:2、9:16 和 16:9 五种画幅比例。模型在生成时直接按指定比例原生构图与布光,无需后期二次裁剪或拉伸。

Grok Imagine Image 2.0 Text to Image 一次调用能生成多张图片吗?

可以。通过在请求中配置 n 参数为 1 至 4 的整数,模型可在单次任务中并行生成多达 4 张遵循同一提示词但构图与细节各异的独立变体,便于对比选优。

Grok Imagine Image 2.0 Text to Image 是如何计费的?

费用按输出图片张数与选定的画质档位精确结算:1K Low 为 8 积分($0.040/张),1K Medium 与 2K Low 均为 12 积分($0.060/张),2K Medium 为 16 积分($0.080/张)。若任务执行异常中断,系统全额自动返还扣除积分。

Grok Imagine Image 2.0 Text to Image 如何理解长文本与复杂场景指令?

模型具备长达 8,000 字符的提示词上下文容量,能够深度解析复杂的多主体关系、背景透视、材质反射与光影层次,在大型叙事场景中准确呈现各项要素。

撰写 Grok Imagine Image 2.0 Text to Image 提示词有哪些实用建议?

建议按照主体特征、空间布局、材质与环境光线、艺术风格的顺序分层撰写,明确主要物体在画面中的相对位置与景深,为高保真渲染提供精准的空间与视觉锚点。