Seedream 4 Text to Image API

bytedance/seedream/v4/text-to-image
1K/2K/4K · 8 aspect ratios · n 1–15

Seedream 4 Text to Image transforms natural prompts into native 4K visuals, supporting bilingual typography and knowledge infographics. It follows spatial physics closely while maintaining balanced compositions across diverse aspect ratios.

Get API Key
Input
Required0/5000
Optional
Optional
Optional
OutputIdle

Generated images will appear here

Estimated cost: —
Continue using

Examples

text-to-image-02-output.jpg

A cinematic documentary photograph inside a historic sulfur bath house in the Abanotubani district of Tbilisi, Georgia, mid-afternoon. In the foreground, an older Georgian attendant in a white cotton wrap and rubber sandals kneels at the rim of a hexagonal stone pool, testing the water with the back of his hand; steam wraps around his forearm. In the middle ground, a younger woman sits on a marble bench with a folded striped towel over her shoulders, her face half-hidden by a drifting steam veil, talking to a friend who is only visible as a silhouette behind a brick arch. In the far background, a high oculus window drops a single hard shaft of daylight onto the wet brick floor, making the steam glow. Correct occlusion between the three figures, every highlight coming from that one oculus, wet stone texture, visible condensation on the brick vault, 35mm lens, natural skin texture, believable hands with five fingers, no text, no logos, no watermark.

text-to-image-01-output.jpg

A scientifically accurate educational infographic of a Persian qanat on cream paper, warm naturalist watercolor linework. Title across the top is only the Chinese phrase printed exactly 坎儿井如何分水. Four panels in a 2x2 grid. Panel 1 caption is the single English word WELL under the Chinese 母井: a gravel mound with a vertical well, a blue dashed water table, and a hanging bucket. Panel 2 caption is the single English word TUNNEL under the Chinese 暗渠: a downhill tunnel with a 1 percent slope and left-to-right flow arrows. Panel 3 caption is the single English word SHAFTS under the Chinese 竖井: four evenly spaced vertical shafts. Panel 4 caption is the single English word OUTLET under the Chinese 龙口: the tunnel emerges at an oasis and splits into two channels labeled with the single words LEFT and RIGHT under 左渠 and 右渠. Use only these English words: WELL, TUNNEL, SHAFTS, OUTLET, LEFT, RIGHT. Each English word stands alone with space around it. Do not add extra letters. Consistent water table and downhill slope. No other text, no watermark.

text-to-image-03-output.jpg

An ultra-wide cinematic landscape photograph of Lake Abbe on the Ethiopia–Djibouti border at first light. In the left third of the foreground, two Afar salt workers in faded cotton wraps and sandals stand beside a low wooden sledge piled with pale salt cakes; one man lifts a slab while the other steadies a long wooden pole. Mid-frame, a shallow pink-white salt crust stretches toward the lake, with a few flamingos standing in a thin sheet of water. On the right, a row of tall pale limestone chimneys rises from the flats, their bases still in blue shadow while their tops catch the first gold sun. Far behind, a hazy volcanic ridge sits under a clear pale sky. Correct relative scale between the workers, the flamingos, the chimneys and the distant ridge, one consistent low sun from the right, fine crystal texture on the salt, visible breath in the cool air, no text, no watermark.

Seedream 4 Text to Image

Seedream 4 Text to Image is a high-fidelity text-to-image model developed by the ByteDance Seed team. Built on a Diffusion Transformer (DiT) and high-compression VAE architecture, it natively outputs up to 4K ultra-high-definition visuals. The model excels in bilingual typography and complex knowledge infographics, powering production-ready advertising and concept design.

Why Choose Seedream 4 Text to Image API?

  • Bilingual Typography & Knowledge-Driven InfographicsNatively renders legible Chinese and English characters, branding slogans, and structured diagrams (such as timelines, flowcharts, and system graphics), overcoming typical character distortion in AI synthesis.

  • Native 4K Ultra-HD Output with Rich TexturesDelivers native ultra-high-definition output up to 4K with no resolution markup, faithfully capturing delicate micro-textures, fabric weaves, cinematic illumination, and natural shadows.

  • Spatial Perspective & Physical CoherenceLeverages advanced multimodal spatial comprehension to respect 3D perspective and depth of field, maintaining accurate proportions and occlusion across complex, multi-element scenes.

  • 1–15 Batch Output & Flexible Aspect RatiosGenerates up to 15 distinct framing and angle variations in a single run across eight standard aspect ratios, accelerating creative explorations at an affordable 5 credits per image.

Parameters

ParameterRequirementDescription
promptRequired

Describe the desired image in 1–5,000 characters. Leading and trailing whitespace is not counted.

sizeOptional

Output aspect ratio. Defaults to 1:1.

Default1:13:44:316:99:163:22:321:9
resolutionOptional

Output resolution preset: 1K, 2K, or 4K. Defaults to 2K.

Default2K1K4K
nOptional

Number of images to generate: integer 1–15. Defaults to 1.

Default123456789101112131415

How to Use

  1. Craft Your Prompt and Text ContentDescribe the scene, subject, aesthetic style, lighting, and camera composition in detail. Enclose any required text in double quotation marks to guide placement and spelling.

  2. Configure Resolution, Ratio, and Image CountSelect 1K, 2K, or 4K resolution, pick from eight aspect ratios such as 1:1, 16:9, or 9:16, and set the image count n between 1 and 15.

  3. Submit Asynchronously and Retrieve VisualsSend your request to the unified endpoint to receive a unique task_id. Poll the status endpoint or configure a callback_url to receive the completed image URLs.

Pricing

Each image costs 5 credits ($0.025). The total is 5 × n credits, where n is the number of generated images.

UsageRateDetails
Standard generation5 credits/image · $0.025/image5 × n credits · $0.025 × n
n = 1 / 5 / 10 / 155 / 25 / 50 / 75 credits$0.025 / $0.125 / $0.250 / $0.375

Best Use Cases

  • Commercial Posters & Bilingual TypographyIdeal for promotional posters, holiday graphics, and social media campaigns requiring legible Chinese and English headlines, brand slogans, and polished typography.

  • Educational Visualization & InfographicsTransforms abstract data, historical timelines, and technical concepts into well-structured, visually appealing diagrams and presentation graphics.

  • E-Commerce Merchandising & Packaging MockupsGenerates studio-grade product presentations, packaging concepts, and hero banners with realistic materials and natural lighting.

  • Concept Art & Storyboard IdeationProduces up to 15 framing and lighting variations per request, helping creative teams rapidly explore concept directions and storyboard iterations.

Pro Tips

  • Enclose Text in Quotation Marks: When rendering specific wording, enclose text in double quotes (for example: banner reading "AI CREATIVE"), specifying font style, alignment, and color hierarchy.
  • Layer Prompts from Subject to Environment: Structure descriptions starting from the main subject, followed by relative spatial positions, surface textures, directional lighting, and overall atmosphere.
  • Prefer Natural Sentences Over Keyword Stacks: Seedream 4 excels at natural language understanding; coherent sentences produce more harmonious compositions than fragmented tag lists.
  • Iterate with Batch Variations: Start with n=4 at 2K resolution for rapid composition exploration, then refine the winning prompt to generate high-resolution 4K production assets.

Notes

  • Text-to-Image Endpoint Scope: This endpoint generates images solely from text prompts without reference image inputs. For image editing, style transfer, or multi-image fusion, use the Seedream 4 Edit endpoint.
  • Prompt Character Limits: Prompts must be between 1 and 5,000 characters after trimming leading and trailing whitespace, and must contain at least one non-whitespace character.
  • Batch Billing & Automatic Refund: Tasks pre-deduct 5 × n credits upon submission. If a task fails due to validation errors or timeouts, all credits are refunded automatically, and partial outputs are refunded proportionally.

Related Models

Seedream 4 Text to Image API frequently asked questions

What is the Seedream 4 Text to Image API?

Seedream 4 Text to Image is a ByteDance Seed model for high-quality text-to-image generation. It generates 1K, 2K, and 4K ultra-high-definition images from natural language prompts with crisp bilingual typography, refined aesthetic lighting, and batch creation of 1–15 images. Built on an efficient Diffusion Transformer (DiT) and high-compression VAE unified architecture, it preserves spatial semantics and physical prompt adherence while maintaining balanced compositions across diverse aspect ratios. You can call it programmatically or try it from the playground above.

Does Seedream 4 Text to Image support rendering Chinese and English text?

Yes. The model features a built-in typography module that accurately renders legible Chinese and English characters, headlines, and slogans directly within generated graphics. Enclosing text in quotation marks and specifying placement hierarchy guides the model to produce sharp, orderly characters.

Can Seedream 4 Text to Image generate multiple images at once?

Yes, you can generate 1 to 15 independent images in a single call. By configuring the integer n parameter from 1 to 15, you receive multiple composition and framing variations for the same prompt, billed only for the images successfully produced.

What resolutions and aspect ratios does Seedream 4 Text to Image support?

It supports three native resolution presets—1K, 2K, and 4K—defaulting to 2K. It also provides eight standard aspect ratio presets: 1:1, 3:4, 4:3, 16:9, 9:16, 3:2, 2:3, and 21:9, defaulting to 1:1.

Does Seedream 4 Text to Image support generating infographics and diagrams?

Yes. The model incorporates extensive world knowledge and visual layout capabilities, understanding conceptual hierarchies and procedural flows to synthesize well-organized diagrams, timelines, and educational infographics directly from prompts.

What is the typical generation latency for Seedream 4 Text to Image?

Benefiting from its DiT backbone and distillation acceleration, Seedream 4 delivers significantly faster inference than its predecessors. Under standard cluster load, end-to-end generation latency for a 2K or 4K image has a median of 5 to 15 seconds. Production applications should poll every 2 to 5 seconds or supply a callback_url.

How are credits handled if a Seedream 4 Text to Image task fails?

Vidgo provides an automated refund guarantee. Tasks pre-deduct 5 credits per image upon submission; if a task fails due to validation errors or timeouts, all credits are returned immediately, and partial output is refunded proportionally.