Image model for text posters and precise reference-image editing
gpt-image-2:official is the image generation and editing model for OpenAI GPT Image 2, suitable for turning text concepts, product photos, and character references into deliverable visual assets. It focuses on text in images, layout composition, and reference-image-guided creation. It can generate posters and infographics from scratch, as well as modify colors, backgrounds, or local objects in existing images, making it suitable for design workflows that require repeated reviews.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Creation mode
Text-to-image generation; reference-image editing; multipart masked local editing
Reference image input
Editing supports a single URL, an array of up to 16 URLs, or local file uploads
Output quantity
1–10 images per request; only 1 image is supported when response_format=b64_json
Canvas control
size=auto or WIDTHxHEIGHT; width and height must be multiples of 16, with the longer side not exceeding 3840 pixels
Pixel range
Total pixels: 655,360–8,294,400; aspect ratio must not exceed 3:1
Quality and files
Common quality options are auto, low, medium, and high; outputs PNG, JPEG, and WebP
Local editing mask
Alpha PNG, no larger than 4MB, with the same dimensions as the first original image; transparent areas can be modified
GPT Image 2 provides image creation and editing capabilities. The quantities, dimensions, formats, and upload rules above correspond to this platform's calling method for this model.
Core Capabilities
Integrate Copy into the Image
Suitable for generating headlines, slogans, explanatory labels, and the main image together, rather than merely drawing a text-free background. When creating, you can specify the main headline position, font-size hierarchy, whitespace, and color palette, guiding a poster or infographic toward the intended visual structure; the final copy, numbers, and reading order should still be checked word by word.
Constrain Creativity with Reference Images
When editing, you can provide product, character, or style references at the same time and explain the role of each image, such as preserving the product appearance, drawing from the background color tone, or adjusting the shooting atmosphere. This is suitable for deriving different creative ideas from existing assets, but reference images are creative constraints and do not mean trademarks, faces, and fine details will remain unchanged pixel by pixel.
Focus Edits on Local Content
Mask editing can focus changes on specified areas, such as replacing objects on a tabletop, adjusting part of an outfit, or adding background elements. Prompts should describe both the complete target image and the content that needs to be retained; clear Alpha boundaries help convey intent, but edge blending should still be checked after the image is generated.
Use Cases
Event Posters and Infographics
Enter the event title, verified information, brand color palette, and layout requirements to generate promotional posters or infographics for proposals. It is recommended to distinguish between text that must be presented accurately and decorative content that can be freely improvised; when delivering, focus on checking dates, prices, labels, and hierarchy, then complete layout proofreading before formal publication.
Product Concepts and Scene Image Editing
Provide a clear product image and specify the packaging, outline, and angle to retain, then try different backgrounds, colors, or lifestyle scenes. Suitable for e-commerce hero visuals, advertising proposals, and packaging presentation drafts; use masks when only local changes are needed, and compare the generated image with the original to check brand identifiers and product details.
Character Design and Multi-Option Proposals
Enter character references, clothing details, expression requirements, and image layout to create character design sheets or creative candidates with the same theme. Use multiple generations when several directions are needed to compare composition and style; for ongoing projects, retain the selected reference images and check the consistency of faces, clothing, and props one by one.
How to choose this model
Choose it when generation and image editing need to be continuous
If a task requires both generating a draft from text and continuing to refine a selected image, GPT Image 2 is suitable as a creative tool in the same workflow. First validate the composition at low or medium quality, then increase quality to inspect details. Existing Image 1.5 projects can use the same copy and reference assets for comparison, focusing on text, layout, and editing results rather than deciding migration solely by version number.
Choose between it and 2.5 variants based on objectives
GPT Image 2.5 Flare emphasizes generation speed, while Sunburst emphasizes high fidelity and fine control; GPT Image 2 is suitable for tasks such as text posters, product creatives, and reference-image editing. Existing qualified templates can continue using this model; if delivery priorities shift to speed or more detailed editing control, compare the relevant variants with real assets without assuming a fixed degree of improvement.
Getting started
Determine whether to use text-to-image or image editing
For generation, provide prompt; for editing, provide both image and modification instructions. Clearly specify the original text, subject-retention requirements, and target aspect ratio.
Specify the model and parameter format
Call the image generation or image editing endpoint, explicitly specifying model=gpt-image-2:official; use auto or WIDTHxHEIGHT for size, and generate one image first before evaluating it. Configure masks, quality, and file format according to this endpoint guide.
Check images and cost records
Read the URL or Base64 image according to the response format, save task_id asynchronously before querying results; check text, reference details, and alpha channels, and record usage according to the current Pricing rules.
Trial recommendation: local background editing with a mask
Input and objective
Keep the product in the original image and modify only the background to a soft gradient; transparent areas of the mask indicate the background that needs editing, while the subject area remains covered.
Acceptance criteria and next steps
Use the multipart upload method specified in the guide for the mask, and verify the image and mask dimensions and alpha channels; do not submit the mask URL as a JSON field.
Usage Limits
Text and structured layouts within images still require human review. Long passages, dense labels, precise alignment, and cross-image character consistency may deviate from requirements; when dates, prices, or chart data are involved, provide accurate content and include character-by-character verification in the delivery process rather than letting the model replace fact-checking.
Masks are not hard-edged selections in traditional image software. When using them, upload the original image and mask together via multipart; do not mix a URL original image with a local mask; the mask must have an Alpha channel. Even if the region is set correctly, edge blending and nearby details may still change.
The canvas must satisfy limits for side lengths, total pixels, and aspect ratio simultaneously; do not consider only one size condition. size=auto is suitable for exploring compositions, while precise delivery should specify pixel values; this model outputs static images and will not automatically perform ChatGPT-style web searches, multi-step orchestration, or video production.
Frequently Asked Questions
Is gpt-image-2:official an independent new model?
It is not another independent foundation model, but a public invocation ID variant of GPT Image 2. Both generation and editing requests should explicitly specify gpt-image-2:official; it shares the same foundational creative positioning as the ID without the suffix, but invocation and billing methods should be distinguished according to the selected model.
Which endpoint should reference images be sent to?
Use /openai/images/generations to create from text; use /openai/images/edits when existing images need modification, passing a single URL, an array of URLs, or uploaded files through image. For multi-image input, explain the subject, style, or scene role of each image.
How do I modify only part of an image?
Use the edits endpoint and place the local original image and mask in the same multipart request. The mask must be an Alpha PNG of the same dimensions, with transparent areas indicating where modifications are allowed; the prompt should also clearly state what to modify and what to preserve, then check the region edges and surrounding details after generation.
Can I return multiple images or Base64 at once?
You can use n to request 1–10 candidate images, making it easier to compare different compositions. response_format can be url or b64_json, but Base64 responses support only a single image. The output file format is separately controlled by output_format, with PNG, JPEG, or WebP available.
How should I handle longer generations and costs?
You can add callback_url so the request first returns task_id, then receive the final result upon completion; the synchronous method returns image data directly. This model is billed based on actual Token usage, and the amount before the request is an estimate. Saving task IDs and deduplicating callbacks helps manage long-running tasks and avoid duplicate processing.