GPT Image 1.5 Edit API

openai/gpt-image-1.5/edit
image_urls · mask_url

GPT Image 1.5 Edit modifies existing images using text instructions and image references, featuring robust facial likeness preservation, optional mask-guided inpainting, and multi-image context blending. It seamlessly integrates new creative changes while preserving original subject identity, natural lighting, and fine background textures.

Get API Key
Input
Required0/1000
Required
Optional
Optional
Optional
OutputIdle

Generated images will appear here

Estimated cost: 2 × 1 = 2 credits · $0.010

2 credits ($0.010) per image, multiplied by n.

Continue using

Examples

edit-03-output.png

Edit only the text on the lower plank: replace "SUMMIT 5 KM" with "SUMMIT CLOSED - SNOW" in the same cream-white hand-brushed capital letters, with the same letter height, stroke texture and slight unevenness, fitted neatly inside the same plank. Leave the upper plank "LAKE TRAIL 2 KM", both arrow shapes, the post, the wood grain, the camera angle and the forest exactly as they are. Then add only a thin layer of fresh snow resting on the top edges of both planks and on top of the post, with a few light snowflakes in the air. No other changes, no extra text, no logos, no watermark.

edit-01-output.png

Turn the creature in this child's crayon drawing into a believable real animal photographed in a sunny backyard garden. Keep its exact design from the drawing: the round purple body, the same large orange spots in the same places, exactly six short stubby legs, two antennae ending in yellow balls, one big eye and the wide crooked smile with two square teeth. Give it realistic soft velvety skin, natural weight and contact shadows on the grass. Eye-level wildlife photography with shallow depth of field, warm afternoon light, a blurred wooden fence and flowers in the background. Friendly and charming, not scary. No paper, no crayon texture, no text, no watermark.

edit-02-output.png

Combine the two reference images. Place the woman from image 1 into the hillside apiary from image 2. She stands beside the second beehive, holding a honeycomb frame she has just lifted out with both gloved hands and inspecting it closely while a few bees fly around. Keep her face, age, silver hair bun, hazel eyes and white beekeeping jacket exactly as in image 1. Keep the painted hive colors, lavender, stone wall, olive trees and hills from image 2. Relight the whole scene to warm golden-hour sunlight coming from the right, with matching shadows on the ground and a soft rim light on her hair, so she looks naturally photographed there. Realistic documentary photograph, three-quarter view, no text, no logos, no watermark.

GPT Image 1.5 Edit

GPT Image 1.5 Edit is OpenAI's multimodal image editing model designed for precision modifications and multi-source visual synthesis. Supporting one or more reference images alongside optional mask inputs, it enables targeted inpainting, style transfers, object addition or removal, and background replacement via natural language instructions. The model excels at facial and identity preservation, supports square, portrait, and landscape sizes, and maintains seamless lighting and texture coherence between edited areas and original surroundings.

Why Choose GPT Image 1.5 Edit?

  • Robust Facial & Identity PreservationDeeply locks subject facial features, contours, and fine details across iterative edits, preventing identity drift during wardrobe, background, or pose changes.

  • Mask-Guided Inpainting & Direct Prompt EditsSupports targeted edits using mask_url for bounded modifications, or executes hair-level precision inpainting directly guided by natural language instructions.

  • Multi-Reference Context BlendingAccepts multiple reference images in image_urls, intelligently combining subject identity, apparel style, props, and backgrounds into a coherent composite.

  • Physically Consistent Lighting & TexturesAccurately simulates ambient lighting, reflections, shadows, and depth of field across edited and untouched regions for seamless integration.

Parameters

ParameterRequirementDescription
promptRequired

Describe the requested image changes using 1–1,000 characters, including at least one non-whitespace character.

image_urlsRequired

Provide at least one publicly accessible HTTP(S) image URL. Enter URLs or upload images.

sizeOptional

Output image dimensions in pixels. Defaults to 1024x1024.

Default1024x10241024x15361536x1024
nOptional

Number of output images: integer 1, 2, 3 or 4. Defaults to 1.

Default1234
mask_urlOptional

Mask image selecting the area to edit. Enter a publicly accessible HTTP(S) URL or upload an image.

How to Use

  1. Provide Reference Image URLsInclude at least one publicly accessible HTTP(S) image URL in the image_urls array as the editing canvas.

  2. Specify Prompt and Optional MaskDescribe desired modifications in prompt (e.g., "replace background with a night beach"); optionally pass a black-and-white mask via mask_url to constrain edits.

  3. Configure Canvas Size and CountSelect desired size (1024x1024, 1024x1536, or 1536x1024) and image count n (1–4), then review the estimated cost.

  4. Submit Request and Retrieve EditsSubmit the request to obtain a task_id, then retrieve finished high-resolution image URLs via status polling or webhook callback.

Pricing

Each image costs 2 credits ($0.010). Total cost is multiplied by n.

UsageRateDetails
Standard generation2 credits/image · $0.010/image2 × n credits · $0.010 × n
n = 1 / 2 / 3 / 42 / 4 / 6 / 8 credits$0.010 / $0.020 / $0.030 / $0.040

Best Use Cases

  • Portrait Retouching & Wardrobe ChangesRefine facial expressions, update clothing materials, and modify hairstyles while preserving complete subject likeness.

  • E-Commerce Product Background ReplacementIsolate merchandise and transport items into upscale lifestyle or outdoor settings without expensive studio reshoots.

  • Object Removal & Flaw InpaintingSeamlessly erase photobombers, watermarks, or unwanted clutter with automatic background texture infilling.

  • Brand Style Adaptation & Visual CompositingSynthesize multiple brand visual assets to produce consistent multi-image campaign variations and aesthetic transfers.

Pro Tips

  • Use a structured 'action + target object + preservation constraints' phrasing, such as 'replace the coffee cup with a book while keeping the person\'s pose, face, and background unchanged.'
  • When performing precise inpainting, pair prompts with a black-and-white mask image (white for edited zones, black for preserved areas) matching the source dimensions.
  • When blending multiple images, explicitly reference source roles in the prompt (e.g., 'the person from image 1 placed in the lighting of image 2').
  • For optimal facial preservation, use clear, evenly lit reference images without heavy obstructions or extreme angles.

Notes

  • image_urls is required and accepts one or more publicly accessible HTTP(S) image URLs to serve as visual references.
  • prompt is required and accepts 1 to 1,000 characters describing target modifications and elements to keep unchanged.
  • mask_url is optional, providing black-and-white spatial guidance for localized inpainting boundaries.
  • size supports 1024x1024, 1024x1536, and 1536x1024 dimensions, with n generating 1 to 4 parallel variations billed at 2 credits ($0.010) per image.

Related Models

GPT Image 1.5 Edit API Frequently Asked Questions

What is the GPT Image 1.5 Edit API?

GPT Image 1.5 Edit is an OpenAI model for precision image editing and inpainting. It combines publicly accessible reference images, natural language edit instructions, and optional masks to generate modified high-fidelity visuals across 1024x1024, 1024x1536, or 1536x1024 dimensions, featuring robust facial likeness retention and multi-image blending. Built on OpenAI's advanced multimodal diffusion editing architecture, it integrates targeted creative modifications while preserving original subject identity, lighting, and ambient composition. You can call it programmatically or try it from the playground above.

How does GPT Image 1.5 Edit maintain facial consistency during edits?

The model leverages specialized identity-preservation representations. By specifying instructions such as 'keep the person\'s face and identity identical', the latent diffusion model locks key facial geometry and landmarks, modifying only target attributes like clothing or background while preventing facial distortion.

Is a mask image required for local edits in GPT Image 1.5 Edit?

No, a mask is optional. The model natively supports text-guided regional editing when you clearly describe which areas to alter and which to keep intact. Supplying a mask_url provides additional spatial bounding for complex or delicate inpainting tasks.

Does GPT Image 1.5 Edit support multiple reference images?

Yes. You can supply multiple image URLs in the image_urls array. The model extracts individual subjects, styles, and props across the provided inputs and composes them into a unified, coherent output guided by your prompt.

Are there additional fees for input images or masks in GPT Image 1.5 Edit?

No. GPT Image 1.5 Edit shares the exact same pricing model as text-to-image generation, with no surcharges for input images or mask URLs. Pricing is strictly 2 credits ($0.010) per output image, and failed tasks are automatically refunded.

Which canvas dimensions does GPT Image 1.5 Edit support?

GPT Image 1.5 Edit supports three native output dimensions: 1024x1024 (1:1 square, default), 1024x1536 (2:3 portrait), and 1536x1024 (3:2 landscape). The model maintains consistent compositional proportion and perspective across all supported aspect ratios.

Can GPT Image 1.5 Edit generate multiple edit variations simultaneously?

Yes. You can generate 1 to 4 distinct edit variations in a single API call by setting the n parameter (defaults to 1). All variations are processed in parallel according to your reference images and prompt, allowing rapid comparison of different edit styles.

How should prompts be written for effective GPT Image 1.5 Edit modifications?

Adopt an explicit edit pattern: define the specific modification, identify the target area, and explicitly list what should be preserved (e.g., 'replace the brick wall with a wooden fence, preserving the cat in the foreground'). Clarity on changes versus invariants produces the most reliable results.