GPT Image 2
OpenAI's latest image generation and editing model with a reasoning thinking mode achieving 99%+ text rendering accuracy, supporting up to 2K resolution, flexible aspect ratios, and multilingual text, deeply integrated into ChatGPT and API.
- Modalities
- Text to Image · Image to Image
- Starting price
- From $0.9375 / call
- Calculator
OpenAI
README
GPT-Image-2 is OpenAI's flagship API model for image generation and editing. It accepts text prompts and high-fidelity image inputs and produces images. OpenAI provides the rolling gpt-image-2 alias and the pinned gpt-image-2-2026-04-21 snapshot; the official documentation does not separately disclose an exact initial launch date, parameter count, or underlying architecture. Developers can use the Image API for direct generation and editing or invoke image generation inside conversational, multi-step Responses API workflows.
Compared with earlier GPT Image models, GPT-Image-2 focuses on stronger instruction following, flexible output dimensions, and automatic high-fidelity handling of reference images. It supports generation from scratch, single- or multi-image composition, masked local edits, and iterative editing. The longest edge can reach 3,840 pixels with a maximum of 8,294,400 total pixels, and quality can be set to low, medium, high, or auto. Outputs can use PNG, JPEG, or WebP with compression controls and preview transparent backgrounds, although OpenAI marks images above 2560×1440 pixels as experimental.
Key Capabilities
- Text-to-Image: Creates photographs, illustrations, concept art, and design assets from natural-language prompts with control over subject, style, composition, and text elements.
- High-Fidelity Editing: Automatically processes every image input at high fidelity, helping preserve people, products, and brand details while changing scenes or styles.
- Multi-Image Composition: Uses multiple reference images together to combine objects, styles, or design elements into a new visual.
- Masked Editing: Restricts changes to a selected region for object replacement, background adjustment, local repair, or canvas extension while preserving other areas.
- Multi-Turn Editing: Iterates on the same visual concept within a Responses API conversation, progressively refining composition, color, copy, and detail.
- Flexible Size and Quality: Accepts custom resolutions within documented constraints, including common square, landscape, portrait, 2K, and experimental 4K sizes, with four quality settings.
- Transparent Background and Formats: Offers preview transparent-background output and supports PNG, JPEG, or WebP, with configurable compression for JPEG and WebP.
Technical Strengths
| Feature | Benefit |
|---|---|
| Automatic High-Fidelity Inputs | Processes all reference images at high fidelity without an input_fidelity setting, helping preserve faces, products, and design details. |
| Flexible Resolution System | Supports custom sizes that meet edge, aspect-ratio, and pixel-count constraints, reducing composition loss from post-generation cropping or reformatting. |
| Unified Generation and Editing | Covers original creation, reference-based transformation, multi-image composition, and local edits in one model and workflow. |
| Conversational Visual Iteration | Responses API context enables multi-turn editing, making clarification and incremental refinement closer to a real design collaboration. |
| Progressive Result Preview | The Image and Responses APIs can stream up to three partial images, reducing uncertainty during generation in interactive products. |
| Configurable Output Controls | Size, quality, format, compression, and background settings let developers trade off fidelity, file size, latency, and delivery requirements. |
Pricing
How the estimate works
Output tokens are calculated from the quality grid, aspect ratio, and total pixel area. Estimated cost equals output tokens multiplied by the price per 1 million output tokens.
- Low · 1024 × 1024: 196 output tokens
- Medium · 1024 × 1536: 1372 output tokens
- High · 1536 × 1024: 5488 output tokens
| Token Type | LinkAI Price | Official Price |
|---|---|---|
| Image · Input | $6 / 1M tokens | $8 / 1M tokens |
| Text · Input | $3.75 / 1M tokens | $5 / 1M tokens |
| Image · Cached input | $1.5 / 1M tokens | $2 / 1M tokens |
| Text · Cached input | $0.9375 / 1M tokens | $1.25 / 1M tokens |
| Output | $22.5 / 1M tokens | $30 / 1M tokens |
| Image · Output | $22.5 / 1M tokens | $30 / 1M tokens |