Seedream 4 Text to Image API

bytedance/seedream/v4/text-to-image
1K/2K/4K · 8 aspect ratios · n 1–15

Seedream 4 Text to Image 将自然语言提示词转化为最高 4K 原生图像,支持中英双语文字排版与知识信息图解。它在严格遵循空间物理规律的同时,保持多画幅构图严谨自然。

获取 API 密钥
Input
必填0/5000
可选
可选
可选
Output待生成

生成的图片将在此处显示

预估费用: —
继续使用

示例

text-to-image-02-output.jpg

A cinematic documentary photograph inside a historic sulfur bath house in the Abanotubani district of Tbilisi, Georgia, mid-afternoon. In the foreground, an older Georgian attendant in a white cotton wrap and rubber sandals kneels at the rim of a hexagonal stone pool, testing the water with the back of his hand; steam wraps around his forearm. In the middle ground, a younger woman sits on a marble bench with a folded striped towel over her shoulders, her face half-hidden by a drifting steam veil, talking to a friend who is only visible as a silhouette behind a brick arch. In the far background, a high oculus window drops a single hard shaft of daylight onto the wet brick floor, making the steam glow. Correct occlusion between the three figures, every highlight coming from that one oculus, wet stone texture, visible condensation on the brick vault, 35mm lens, natural skin texture, believable hands with five fingers, no text, no logos, no watermark.

text-to-image-01-output.jpg

A scientifically accurate educational infographic of a Persian qanat on cream paper, warm naturalist watercolor linework. Title across the top is only the Chinese phrase printed exactly 坎儿井如何分水. Four panels in a 2x2 grid. Panel 1 caption is the single English word WELL under the Chinese 母井: a gravel mound with a vertical well, a blue dashed water table, and a hanging bucket. Panel 2 caption is the single English word TUNNEL under the Chinese 暗渠: a downhill tunnel with a 1 percent slope and left-to-right flow arrows. Panel 3 caption is the single English word SHAFTS under the Chinese 竖井: four evenly spaced vertical shafts. Panel 4 caption is the single English word OUTLET under the Chinese 龙口: the tunnel emerges at an oasis and splits into two channels labeled with the single words LEFT and RIGHT under 左渠 and 右渠. Use only these English words: WELL, TUNNEL, SHAFTS, OUTLET, LEFT, RIGHT. Each English word stands alone with space around it. Do not add extra letters. Consistent water table and downhill slope. No other text, no watermark.

text-to-image-03-output.jpg

An ultra-wide cinematic landscape photograph of Lake Abbe on the Ethiopia–Djibouti border at first light. In the left third of the foreground, two Afar salt workers in faded cotton wraps and sandals stand beside a low wooden sledge piled with pale salt cakes; one man lifts a slab while the other steadies a long wooden pole. Mid-frame, a shallow pink-white salt crust stretches toward the lake, with a few flamingos standing in a thin sheet of water. On the right, a row of tall pale limestone chimneys rises from the flats, their bases still in blue shadow while their tops catch the first gold sun. Far behind, a hazy volcanic ridge sits under a clear pale sky. Correct relative scale between the workers, the flamingos, the chimneys and the distant ridge, one consistent low sun from the right, fine crystal texture on the salt, visible breath in the cool air, no text, no watermark.

Seedream 4 Text to Image

Seedream 4 Text to Image 是字节跳动 Seed 团队研发的高保真文生图大模型。模型基于 Diffusion Transformer (DiT) 与高压缩比 VAE 架构,原生支持最高 4K 超高清分辨率直出。它在中英双语文字排版与复杂知识图表生成上表现优异,单次请求支持并发生成多张创意成品,适用于海报制作与商业视觉设计。

为什么选择 Seedream 4 Text to Image API?

  • 中英双语文字与密集图表排版原生支持在生成图像中精准呈现清晰工整的中英双文字符、品牌标语与专业知识图表(如流程图、时间线与架构图表),解决传统 AI 绘图文本扭曲与乱码难题。

  • 原生 4K 超高清直出与细腻质感最高支持 4K 原生超高清分辨率直出且无需额外加价,细腻呈现真实皮肤毛孔微结构、织物纹理、电影级自然光照与真实阴影投射。

  • 复杂透视与物理空间规律推理依托深层多模态空间理解能力,严格遵循透视法则与 3D 景深层次,准确呈现多主体间的大小比例与前后遮挡关系,确保画面纵深自然协调。

  • 单次 1–15 张并发发散与灵活画幅单次请求即可并发生成最多 15 张不同视角与构图方案的独立画面变体,原生适配 8 种画幅比例,配合单张 5 积分平价计费,全面提升创意选型效率。

参数说明

参数要求说明
prompt必填

用 1–5,000 个字符描述想要生成的图片,首尾空白不计入长度。

size可选

输出画幅比例,默认 1:1。

默认值1:13:44:316:99:163:22:321:9
resolution可选

输出分辨率档位,可选 1K、2K、4K,默认 2K。

默认值2K1K4K
n可选

生成图片数量,取值为整数 1–15,默认 1。

默认值123456789101112131415

使用教程

  1. 撰写提示词与指定画面文字使用详尽的中英文描述画面主体、艺术风格、光影质感与构图机位;若需要在画面中呈现特定文字,请使用半角双引号明确标注文案内容及排版位置。

  2. 配置分辨率、画幅与输出张数根据应用媒介选择 1K、2K、4K 清晰度档位,指定 16:9、9:16、1:1 等 8 种常用比例,并通过 n 参数设定 1 至 15 张期望生成的图片数量。

  3. 异步提交任务并获取高清图片向统一端点发送请求并获取唯一 task_id,通过轮询查询接口或在顶层配置 callback_url 接收终态通知,任务完成后直接读取并下载高清图片 URL。

价格

每张图片 5 credits($0.025),总价按 n 计算:5 × n credits。

用量价格说明
标准生成5 credits/image · $0.025/image5 × n credits · $0.025 × n
n = 1 / 5 / 10 / 155 / 25 / 50 / 75 credits$0.025 / $0.125 / $0.250 / $0.375

适用场景

  • 商业海报与双语文字排版适用于需要融合端正中英文字题字、品牌 Slogan 与排版艺术的商业宣传海报、节日贺卡与社交媒体营销物料。

  • 知识科普与信息图解呈现将抽象数据、历史时间线、技术原理与业务流程转化为层次清晰、图文并茂的专业信息图表与演示幻灯片素材。

  • 电商商品陈列与包装设计结合逼真材质渲染与自然光源模拟,为消费品快速生成具备商业质感的展示场景、包装概念图与广告主视觉。

  • 影视前期概念与分镜批量发散利用 1–15 张多图批量发散能力,在统一艺术设定下快速探索多种景别、机位与氛围方案,加速项目前期方案定稿。

使用技巧

  • 画面文字双引号规范:若需在画面中精准渲染文字,请在提示词中用半角双引号指明文案(例如:海报上方写着 "AI CREATIVE"),并说明字体的风格、排版层次与颜色。
  • 空间层次递进描述:按照“主要主体 + 相对空间方位 + 材质纹理细节 + 动态光源环境 + 全局艺术基调”的顺序组织提示词,能够更好地引导模型构建合理的 3D 纵深与光影投射。
  • 使用自然描述避免关键词堆叠:Seedream 4 具备出色的自然语言理解底座,使用完整连贯的语句描述画面比散乱罗列修饰词能够获得更加严谨自然的构图效果。
  • 阶梯式批量发散与定稿:方案探索初期建议以默认 2K 尺寸设置 n 为 4 进行快速构图发散选型;确定理想构图与美学风格后,再针对性优化提示词生成 4K 最终成品。

注意事项

  • 纯文本生图专属端点:本端点专用于通过纯文本提示词生成图像,不接收图片素材;如需结合参考图进行编辑、风格迁移或多图融合,请调用 Seedream 4 Edit 端点。
  • 提示词长度规范:提示词去除首尾空白字符后的长度需在 1 至 5,000 字符之间,包含至少一个非空白字符。
  • 批量计费与退款保障:任务提交时按 5 × n 积分预扣除;若因校验未通过、网络异常或超时导致生成失败,预扣积分将全额自动退还;若实际成功出图少于 n 张,系统按未出图差额自动退费。

相关模型

Seedream 4 Text to Image API 常见问题

Seedream 4 Text to Image API 是什么?

Seedream 4 Text to Image 是字节跳动 Seed 团队用于高质量文本生成图像的多模态大模型。它根据自然语言提示词生成 1K、2K 与 4K 超高清图像,具备行业领先的中英双语文字排版渲染能力、细腻真实的美学光影质感,并支持单次 1–15 张批量生成。基于高效的 Diffusion Transformer (DiT) 与高压缩比 VAE 统一架构,它在严格遵循提示词空间几何与物理规律的同时,保持多画幅构图严谨自然。你可以通过 API 进行程序化调用,也可以在上方体验区直接在线试用。

Seedream 4 Text to Image 支持在画面中渲染中文和英文文字吗?

支持。模型内置强大的文字渲染模块,原生支持在生成图像中精准呈现清晰工整的中英双文字符、品牌标语与海报题字。在提示词中用半角双引号明确标注文案并在提示词中指明排版位置与字号层次,即可生成端正锐利的字迹。

Seedream 4 Text to Image 一次能批量生成多张图片吗?

支持单次请求生成 1 至 15 张独立图片。通过在请求参数中配置 n(取值整数 1–15),即可针对同一提示词并发获得多种构图与视角的发散变体,计费按实际成功输出张数结算。

Seedream 4 Text to Image 支持哪些原生分辨率与画幅比例?

支持 1K、2K 与 4K 三种原生清晰度档位,默认档位为 2K;同时支持 1:1、3:4、4:3、16:9、9:16、3:2、2:3、21:9 等 8 种主流比例预设,默认画幅为 1:1。

Seedream 4 Text to Image 支持生成信息图表和知识图解吗?

支持。模型融合了广博的知识库与视觉设计排版能力,能够理解复杂学科概念、流程逻辑与多级层级结构,根据提示词直接生成布局规范、视觉清晰的科普信息图解、流程图与对比示意图。

Seedream 4 Text to Image 生成耗时通常在什么范围?

得益于 DiT 骨干与对抗蒸馏加速技术,Seedream 4 的推理效率相比前代实现了显著提升。在标准并发与网络负载下,单张 2K 或 4K 图像的端到端生成耗时中位数约为 5 至 15 秒;批量生成多张图片时耗时会略有增加,生产环境建议配置 2 至 5 秒轮询间隔或在请求中传入 callback_url 异步接收结果。

Seedream 4 Text to Image 任务生成失败时如何结算积分?

系统具备全自动退款保障机制。任务提交时会按 5 × n 积分进行预扣除;若任务因参数校验不通过、网络异常或系统超时导致生成失败,预扣积分将即时全额退还;若实际成功生成图片少于 n 张,系统会自动按未出图差额退还积分。