Grok Imagine Image Quality Text to Image API

xai/grok-imagine-image-quality/text-to-image
1K / 2K · 5 种比例

Grok Imagine Image Quality Text to Image 将纯文本提示词转化为最高 2K 分辨率的高保真图像,支持多语言排版文字渲染、微小细节精准把控,以及 5 种主流构图比例。它能够严格遵循光学镜头规格与复杂布光指令,在保留场景自然质感与风格连贯性的同时,呈现细腻逼真的光影细节与真实物理纹理。

获取 API Key
输入
973
输出已就绪
1 × 11 = 11 积分 · $0.055

继续使用

示例

text-01.png

A photorealistic documentary photograph inside a basalt sea cave at low tide. One adult woman geologist with short dark hair, a weathered ochre field jacket and a small canvas backpack studies the columnar rock wall in three-quarter profile. Her right hand holds a small warm-white flashlight aimed at a nearby mineral seam; her left hand rests naturally on a closed field notebook at her waist. Cool blue daylight from the cave entrance on the left meets the restrained amber flashlight pool on the right. Believable wet basalt microtexture, small tidal reflections, natural skin pores and damp hair, fine sea spray in the distant opening. Medium-wide eye-level 35mm composition, person on the right third, coherent perspective and natural dynamic range, face and near rock in focus, distant ocean softer. Quiet field research, no theatrical posing, no additional people, no writing. No advertisement, brand, logo, fox, product display, promotional graphics, or watermark.

text-02.jpg

An intimate wildlife photograph of exactly one Japanese macaque resting beside the exposed roots of a cedar tree in a snowy mountain forest. The macaque is seated in a natural relaxed posture, three-quarter view, with both hands resting loosely in its lap and a small visible breath cloud in the cold air. Fine silver-brown fur with individual snow crystals, warm pink face, alert dark eyes, anatomically believable hands. Soft overcast morning light, cool white snow, warm natural fur, distant cedar trunks gently out of focus. Eye-level 85mm lens, shallow but credible depth of field, entire face sharply focused. Vertical composition with uncluttered negative space above, a candid moment without human objects or clothing. No text. No advertisement, brand, logo, fox, product display, promotional graphics, or watermark.

text-03.jpg

A cinematic yet realistic street-level architectural photograph of a small community library at night. Through a broad clear window, exactly two adult readers sit separately at a long wooden reading table beneath warm pendant lamps; shelves recede with consistent perspective. Outside is a quiet dry sidewalk and a deep blue evening sky, no rain. Beside the entrance, one simple illuminated rectangular wayfinding sign displays exactly two lines: "夜读" on the first line and "NIGHT READING" on the second. Both lines must be correctly spelled, sharp, evenly spaced, and easily readable. Frame the sign large enough to read, in the near left third; the window and readers fill the rest. Calm amber interior versus blue exterior, believable glass reflections that do not obscure people or lettering. No other legible text, shop branding, banners, posters or sales context. No advertisement, brand, logo, fox, product display, promotional graphics, or watermark.

Grok Imagine Image Quality Text to Image 概览

Grok Imagine Image Quality Text to Image 是 xAI 针对高逼真度创作需求打造的高保真文本生成图像模型。它以极佳的照片真实感、精准的复杂提示词遵循度、原生多语言清晰文字排版与光学级质感为核心优势,提供 1K 与 2K 两种分辨率档位,原生支持 1:1、2:3、3:2、9:16 与 16:9 五种画幅比例。

为什么选择 Grok Imagine Image Quality Text to Image?

  • 照片级真实感与细腻光影专为高保真画质调优,呈现自然的皮肤纹理、微瑕细节与逼真的环境光影交互,营造具有真实呼吸感的摄影级画面。

  • 多语言文字排版锐利渲染突破传统图像模型难以生成清晰文字的瓶颈,高精度渲染短句与多语言标题,让商业海报与包装设计上的文字工整清晰。

  • 光学级专业提示词遵循深度理解相机镜头(如 35mm、85mm)、光圈景深、布光方案及复杂构图指令,准确转化多层次视觉创意。

  • 1K 与 2K 分辨率及多画幅适配提供 1K 快速迭代与 2K 旗舰画质双档位,原生支持 1:1、2:3、3:2、9:16 与 16:9 五种比例,灵活适配各类媒介场景。

  • 透明阶梯计费与全额退还保障按输出分辨率与张数精确核算积分,若生成任务异常中断或失败,系统自动全额返还扣除积分,保障每一分预算。

参数

参数要求说明
prompt必填

必填非空字符串;上游未声明长度上限。

aspect_ratio可选

支持 1:1、2:3、3:2、9:16、16:9,默认 1:1。

默认1:12:33:29:1616:9
resolution可选

选择 1K 或 2K,默认 1K。

默认1K2K
n可选

输出图片数量,整数且至少为 1,默认 1;上游未声明最大值。

默认1

使用教程

  1. 编写结构化提示词按照“主体特征 + 构图布光 + 镜头规格 + 材质细节 + 排版文字”层层描述画面,目标文字可用半角双引号括起。

  2. 选择目标分辨率根据输出用途选择 1K 或 2K 分辨率,1K 适合快速构思与预览,2K 提供更高锐度与细节呈现。

  3. 指定画面比例与张数选择适配目标展示平台的宽高比(1:1、2:3、3:2、9:16 或 16:9),并通过 n 设置单次生成的图片张数。

  4. 发起异步生成任务提交请求并获取唯一 task_id,由后台高保真图像集群异步执行渲染处理。

  5. 轮询获取成品图片轮询任务状态接口至 finished 状态,读取生成结果中的高清图像下载链接。

价格

总积分 = 输出数量 × 分辨率单价 + 参考图数量 × 2。参考图费用仅用于编辑,每次请求计一次,不乘输出数量。1 积分 = $0.005。

用量费率详情
1K 输出8 积分 / 输出图片$0.040 / image · Official $0.050 · 20%
2K 输出11 积分 / 输出图片$0.055 / image · Official $0.070 · 21%

最佳应用场景

  • 商业静物与美食摄影精准模拟专业影棚布光与光学景深,生成色彩饱满、质感诱人的高端菜品与商业产品展示图像。

  • 品牌营销海报与图文排版利用卓越的文字排版渲染能力,直接合成包含清晰品牌标语、促销信息与多语言文案的宣发物料。

  • 影视概念艺术与分镜场景将剧本描述与宏大世界观转化为光影考究、情绪饱满的电影级概念静帧与分镜草图。

  • 社交媒体与多端创意物料自由切换 9:16 竖屏短视频封面、16:9 横版头图或 1:1 帖子配图,实现多平台物料的高效批量输出。

专家技巧

  • 加入具体的光学与相机参数(例如“shot on 85mm lens, f/1.4, shallow depth of field, warm morning light”)可显著强化画面的景深透视与专业质感。
  • 需要在画面中准确渲染英文字符或短标语时,使用半角双引号明确框定文字(如 "ARTISAN ROAST"),能获得最高精度的字体排版效果。
  • 若追求真实胶片或抓拍质感,可添加“candid style, subtle film grain, natural skin texture and imperfections”,避免塑料感过重的过度平滑效果。
  • 单次请求可通过设置 n 参数同时生成多张变体,在保持相同布光与构图 prompt 的前提下快速筛选出最具张力的构图方案。

注意事项

  • prompt 为必填非空字符串参数,支持详尽描述画面主体、空间关系与光影技术规格。
  • 分辨率支持 1K(默认)与 2K,画面比例支持 1:1(默认)、2:3、3:2、9:16 与 16:9 五种规格。
  • 任务执行采用异步处理机制,提交后立即返回 task_id,可通过轮询或配置 callback_url 接收最终图片结果。

相关模型推荐

Grok Imagine Image Quality Text to Image API 常见问题

Grok Imagine Image Quality Text to Image API 是什么?

Grok Imagine Image Quality Text to Image 是 xAI 研发的高保真文本生成图像模型。它根据文本提示词直接生成 1K 与 2K 分辨率的高保真图像,具备顶尖的照片真实感、精准的光学布光控制、锐利的多语言排版文字渲染与 5 种构图画幅支持。基于多模态高保真深度生成架构,它在严格遵循复杂技术镜头指令与场景构图的同时,真实呈现细腻的物理质感、自然微瑕与风格一致性。你可以通过 API 进行程序化调用,也可以在上方体验区直接在线试用。

Grok Imagine Image Quality Text to Image 支持在画面中渲染清晰文字吗?

支持。文字渲染是该模型质量层的核心强项之一,在生成商品包装、宣发海报、街景标牌等包含文字的内容时,能够输出字迹工整锐利的短语与多语言排版。在提示词中建议使用英文半角双引号括起目标文字(例如 "SUMMER SALE"),并明确指定文字在画面中的位置与排版样式。

Grok Imagine Image Quality Text to Image 的 1K 和 2K 分辨率该如何选择?

1K 分辨率输出速度更快且单张消耗积分更低(8 积分/张),非常适合快速探索创意草案、社交媒体常规配图或批量初选;2K 分辨率(11 积分/张)能呈现更精细的发丝纹理、微小材质质感与更高锐度,更适合商业印刷品、高清大图宣发或最终交付成片。

如何通过提示词提升 Grok Imagine Image Quality Text to Image 的照片真实感?

建议在提示词中使用具体的光学和相机描述,如指定焦段(35mm、85mm)、光圈(f/1.4、f/2.8)、景深、胶片颗粒度以及真实布光方向(如柔和侧逆光、影棚柔光箱)。同时可搭配 "candid photography"、"natural imperfections" 等关键词,引导模型生成更具自然呼吸感与微瑕细节的真实画面。

Grok Imagine Image Quality Text to Image 支持哪些画面构图比例?

模型原生支持 5 种常用画面比例,包括 1:1(方形,默认)、2:3(竖版人像)、3:2(经典横幅摄影)、9:16(移动端竖屏短视频画幅)和 16:9(宽屏影视构图)。你可以在提交任务时通过 aspect_ratio 参数直接指定,模型会根据目标画幅自动优化画面布局与透视关系。

Grok Imagine Image Quality Text to Image 后续会迁移至哪个版本?

根据 xAI 官方技术规划,Grok Imagine Image Quality 计划于 2026 年 11 月 2 日完成服务过渡,推荐后续开发与业务流程平滑迁移至新一代 Grok Imagine Image 2.0。新版本在保持卓越文字排版能力的同时,进一步强化了多层次提示词遵循度与生成效率。

Grok Imagine Image Quality Text to Image 生成的图片可以商用吗?

可以。根据 xAI 服务条款,调用该模型生成的图片成果物所有权归用户所有,可合法用于商业广告、品牌宣发、数字内容创作及周边衍生品开发。使用者需自行确保输入提示词合规,并遵循平台使用条款与相关版权法规。