GPT Image 1.5 Text to Image API

openai/gpt-image-1.5/text-to-image
1024x1024 · 1024x1536 · 1536x1024

GPT Image 1.5 Text to Image transforms text prompts into high-fidelity visuals, featuring precise instruction adherence, accurate in-image typography, and photorealistic material textures. It preserves complex spatial layouts and lighting coherence while rendering vibrant scenes across square, portrait, and landscape dimensions.

Get API Key
Input
Required0/1000
Optional
Optional
OutputIdle

Generated images will appear here

Estimated cost: 2 × 1 = 2 credits · $0.010

2 credits ($0.010) per image, multiplied by n.

Continue using

Examples

text-to-image-01-output.png

A four-panel comic page in a clean 2x2 grid with thin white gutters, warm hand-inked storybook style with soft watercolor fills. The same character appears in every panel: a small garden snail mail carrier with a teal spiral shell, a tiny red cap and a brown leather satchel. Panel 1: at a mushroom post office, the snail proudly takes a letter stamped EXPRESS; speech bubble: "Express delivery? Leave it to me!" Panel 2: light rain on clover leaves, the snail sliding on under a leaf umbrella; bubble: "Just a little drizzle..." Panel 3: the snail inching across a garden hose like a bridge at sunset; bubble: "Almost there!" Panel 4: winter snow on the ground at a frog's round wooden door, the frog in a scarf reading the letter; frog's bubble: "My birthday party... was last summer." Crisp, correctly spelled lettering in every bubble, consistent character design across panels, gentle humor, no logos, no watermark.

text-to-image-02-output.png

Photorealistic documentary photograph inside a small-town radio studio at 2 a.m. A middle-aged host with a short gray beard and headphones around his neck leans toward a vintage broadcast microphone, one hand resting on a worn analog mixing console with glowing amber VU meters. Above the studio window a red illuminated sign reads "ON AIR". Taped to the wall beside him is a handwritten index card with four lines: "NIGHT OWL SHOW", "2:00 Call-in: Rosa, Maple St.", "2:15 Fog report", "2:30 Lullaby hour". A shelf of cassette tapes whose spines carry only handwritten dates such as "MAR 94" and "OCT 96", a steaming enamel mug, a desk lamp casting warm tungsten light, cool blue streetlight through the window. Shallow depth of field, realistic skin texture, subtle film grain, every piece of text sharp and correctly spelled, no band or artist names, no brand names, no logos, no watermark.

text-to-image-03-output.png

An educational cross-section illustration of a beaver lodge in a forest pond, drawn like a clean natural-history textbook plate with soft watercolor and fine ink lines on off-white paper. The dome of sticks and mud rises above the water, and a cutaway reveals the inside. One adult beaver rests in a dry chamber above the waterline while a second beaver swims up through a submerged tunnel. Title at the top: "Inside a Beaver Lodge". Exactly five labels in neat black sans-serif text, each with a thin leader line pointing to the correct part: "Living Chamber", "Underwater Entrance", "Air Vent", "Winter Food Cache", "Mud and Stick Walls". Show a clear water surface line and a pile of leafy branches stored on the pond floor as the food cache. Accurate beaver anatomy, balanced layout with generous margins, no other text, no logos, no watermark.

GPT Image 1.5 Text to Image

GPT Image 1.5 Text to Image is OpenAI's flagship multimodal text-to-image generation model built on advanced transformer diffusion architecture. Featuring superior semantic understanding and multi-step prompt reasoning, it delivers native in-image typography, photorealistic lighting, and fine micro-texture rendering with generation speeds up to 4x faster than prior generations. The model natively outputs 1024x1024 square, 1024x1536 portrait, and 1536x1024 landscape dimensions, and generates up to 4 consistent variations per request for professional design, advertising, and digital production.

Why Choose GPT Image 1.5 Text to Image?

  • Deep Multimodal Reasoning & Multi-Subject LayoutsPowered by advanced multimodal reasoning, the model parses intricate multi-layered prompts to render complex scenes and multi-subject relationships with exceptional fidelity.

  • Native In-Image Typography & Clear SignageRenders crisp, legible English text on signage, packaging, labels, and posters without common artifacting or spelling distortions.

  • Photorealistic Micro-Textures & Ambient LightingFaithfully renders skin details, fabric textures, liquid reflections, and ambient diffuse lighting without plastic artifacts.

  • Parallel 1–4 Image Variation GenerationGenerates 1 to 4 distinct visual variations in a single request, accelerating creative concepting and design exploration.

Parameters

ParameterRequirementDescription
promptRequired

Describe the image to generate using 1–1,000 characters, including at least one non-whitespace character.

sizeOptional

Output image dimensions in pixels. Defaults to 1024x1024.

Default1024x10241024x15361536x1024
nOptional

Number of output images: integer 1, 2, 3 or 4. Defaults to 1.

Default1234

How to Use

  1. Draft a Structured PromptDescribe the primary subject, visual style, spatial environment, and lighting (up to 1,000 characters). Enclose target in-image text in quotes.

  2. Select Canvas DimensionsSpecify the size parameter based on your use case: 1024x1024 (square), 1024x1536 (portrait), or 1536x1024 (landscape).

  3. Configure Image CountSet the n parameter from 1 to 4 images per generation (defaults to 1), with billing scaling linearly per image.

  4. Submit and Retrieve VisualsPost the request to receive a unique task_id, then track status via polling or webhook callback to download your generated images.

Pricing

Each image costs 2 credits ($0.010). Total cost is multiplied by n.

UsageRateDetails
Standard generation2 credits/image · $0.010/image2 × n credits · $0.010 × n
n = 1 / 2 / 3 / 42 / 4 / 6 / 8 credits$0.010 / $0.020 / $0.030 / $0.040

Best Use Cases

  • Brand Advertising & Social Media VisualsCreate striking promotional banners and social media campaign visuals with custom signage text and vivid lighting.

  • E-Commerce Product ShowcasesRender realistic commercial product staging with accurate depth of field and ambient lighting for merchandise.

  • Poster Headlines & Package DesignDesign creative event posters, typographic flyers, and packaging mockups with clear embedded textual labels.

  • Concept Art & Storybook IllustrationsTranslate creative story ideas into expressive digital art with cinematic mood, rich textures, and stylistic consistency.

Pro Tips

  • Structure prompts logically: define subject characteristics, composition, lighting ambiance, material textures, and any target text in order.
  • When rendering in-image text, enclose the desired English phrase in quotes (e.g., a coffee cup with the text "Morning Brew") and keep phrases concise.
  • Choose 1536x1024 for landscape banners and desktop wallpapers, 1024x1536 for mobile feeds and stories, and 1024x1024 for square avatars or merchandise.
  • Set n to 4 during early concept exploration to evaluate four distinct variations in a single generation round.

Notes

  • prompt is required and accepts 1 to 1,000 characters to guide composition, lighting, and style.
  • size supports 1024x1024 (square), 1024x1536 (portrait), and 1536x1024 (landscape), defaulting to 1024x1024.
  • n sets the number of parallel output variations from 1 to 4 (defaults to 1), with billing calculated as 2 credits ($0.010) multiplied by n.
  • This endpoint generates brand-new visuals directly from text prompts; for image editing, inpainting, or reference-based variations, select the GPT Image 1.5 Edit endpoint.

Related Models

GPT Image 1.5 Text to Image API Frequently Asked Questions

What is the GPT Image 1.5 Text to Image API?

GPT Image 1.5 Text to Image is an OpenAI model for image generation from text prompts. It creates high-fidelity visuals across 1024x1024, 1024x1536, or 1536x1024 dimensions with precise instruction following, photorealistic material textures, and clear in-image typography. Built on OpenAI's advanced multimodal diffusion architecture, it preserves spatial perspective and lighting coherence while rendering complex scenes and text. You can call it programmatically or try it from the playground above.

Does GPT Image 1.5 Text to Image support rendering in-image text?

Yes. The model features enhanced typography capabilities, allowing it to render legible English words on signage, posters, and product labels. Enclosing your target text in double quotes within the prompt (such as a storefront with "Coffee Lab") ensures clear font presentation and accurate spelling.

Which image dimensions does GPT Image 1.5 Text to Image support?

The model supports three standard dimensions: 1024x1024 (1:1 square, default), 1024x1536 (2:3 portrait), and 1536x1024 (3:2 landscape). Generations are composed natively at the chosen aspect ratio without cropping.

How many image variations can GPT Image 1.5 Text to Image generate in one request?

You can generate 1 to 4 images per request using the n parameter (defaults to 1). All variations are processed in parallel based on the same prompt, allowing fast comparison of different compositions and color palettes.

How does GPT Image 1.5 Text to Image calculate credits per generation?

Billing is strictly per generated image at 2 credits ($0.010) per image. Total cost equals 2 × n credits ($0.010 × n), with identical rates across all supported dimensions. If a task fails due to a system error, all deducted credits are refunded automatically.

How should prompts be structured for GPT Image 1.5 Text to Image?

Organize prompts hierarchically: detail the core subject, camera framing (e.g., close-up, wide-angle), ambient lighting (e.g., golden hour side lighting), and surface textures. Concrete physical descriptors yield superior results compared to generic quality buzzwords.

Can images generated by GPT Image 1.5 Text to Image be used commercially?

Yes. Subject to OpenAI's standard content policies, users retain full commercial rights to images generated through the API, allowing deployment in digital marketing, commercial packaging, website UI design, and published merchandise.

When should I choose GPT Image 1.5 Text to Image over GPT Image 1.5 Edit?

Use GPT Image 1.5 Text to Image when generating brand-new visuals entirely from written concepts. If you already have an existing image and wish to perform inpainting, background modification, or subject edits, use the GPT Image 1.5 Edit endpoint instead.