OpenAI Launches ChatGPT Images 2.5: 50% Lower Latency, Sketch, Precision Editing, GPT-Image-2.5 Flare, Sunburst, and 2× API Pricing
- 4 minutes ago
- 4 min read

OpenAI has released ChatGPT Images 2.5, a new image-generation and editing model positioned as its state-of-the-art visual system, with a product update that changes both the consumer creation loop inside ChatGPT and the economics of image generation through the API.
The release combines sharper visual detail, stronger reference-photo fidelity, more precise localized edits and better consistency across repeated editing turns with an OpenAI claim of up to 50% lower image-generation latency than Images 2.0.
ChatGPT also gains Sketch, direct comments on images and format templates, while developers receive two API variants: GPT-Image-2.5 Flare for the default high-volume path and GPT-Image-2.5 Sunburst for premium workflows that prioritize tighter control over editing.
The scale behind the update is already substantial: OpenAI says users create more than 3 billion images every week across ChatGPT Images and the GPT-Image models in the API.
The other material change is price: every listed image/text token price for Flare and Sunburst is exactly 2× the corresponding GPT-Image-2 rate, creating a clear speed-versus-cost trade-off for production workloads.
··········
CHATGPT IMAGES 2.5 CHANGES THE EDITING LOOP.
Fidelity, local edits and multi-turn consistency are the core model-level changes, while latency is the headline performance claim.
Images 2.5 is designed to keep a generated asset anchored to the original reference more reliably, preserving recognizable subjects while changing setting, style, composition or individual visual elements.
For localized editing, the practical goal is to alter only what the user requested while keeping the subject, composition, surrounding details and brand treatment stable rather than forcing a broader regeneration.
That behavior extends across multiple turns: earlier edits are more likely to remain intact as new instructions are layered onto the same asset, reducing the quality drift that can make iterative production workflows expensive in both time and retries.
OpenAI also says the model handles complex visual instructions more coherently, improves the accuracy of real-world information embedded in images, and is better at layouts that require transparent backgrounds.
........
Capability | Images 2.0 / earlier workflow | ChatGPT Images 2.5 |
Generation latency | Baseline | Up to 50% lower |
Reference-photo fidelity | Earlier generation | Improved subject preservation |
Localized editing | Higher risk of unintended changes | More focused edits with surrounding details preserved |
Multi-turn consistency | More drift across repeated edits | Better retention of earlier changes |
Creation controls | Primarily prompt-led | Sketch, image comments, and templates |
Complex layouts | Earlier capability | Improved handling, including transparent backgrounds |
........
··········
SKETCH, COMMENTS, AND TEMPLATES MOVE IMAGE CREATION CLOSER TO AN EDITOR.
Visual instructions now sit alongside text prompts as first-class controls for creating and revising images.
Sketch lets a user draw directly inside ChatGPT and pass that drawing to the model as a visual guide, which is useful when spatial relationships, silhouettes or layout intent are easier to communicate graphically than through a long prompt.
The feature can be invoked with @Sketch, while image comments provide a second form of localized control by letting the user point to a specific area and request a focused modification.
Templates such as Poster and Merch reduce the blank-canvas problem by giving the model a predefined creative format before the user supplies content, style and design constraints.
Taken together, these controls shift the workflow from repeated prompt-and-regenerate cycles toward something closer to an interactive editor: establish structure, generate, mark a region, revise it and preserve the rest of the asset.
That distinction matters most for marketing, retail, presentation and brand workflows, where the value of a model is not only whether it can make a strong first image but whether it can maintain identity and layout through several controlled revisions.
··········
FLARE AND SUNBURST SPLIT THE API INTO SPEED AND PRECISION TIERS.
Flare is the default API model for most applications; Sunburst targets premium work that benefits from tighter edit control.
GPT-Image-2.5 Flare carries the Images 2.5 quality, editing and speed improvements into the API and is positioned as the default option for creator tools, social content, product experiences, visual search, rapid prototyping and high-volume generation.
OpenAI specifically describes Flare as producing higher-quality images than GPT-Image-2 at 50% lower latency, making throughput one of the main reasons to move a latency-sensitive workload to the new generation.
GPT-Image-2.5 Sunburst targets premium visual workflows where detailed control across edits matters more than minimum generation time, including production-ready campaign assets and polished product imagery.
Despite those different workload positions, the two 2.5 models currently share the same published per-token prices.
........
API price metric | GPT-Image-2 | GPT-Image-2.5 Flare / Sunburst |
Image input / 1M tokens | $4 | $8 — 2× |
Cached image input / 1M tokens | $1 | $2 — 2× |
Image output / 1M tokens | $15 | $30 — 2× |
Text input / 1M tokens | $2.50 | $5 — 2× |
Cached text input / 1M tokens | $0.625 | $1.25 — 2× |
........
··········
THE 2× PRICE AND 50% LATENCY CLAIM CREATE A DIFFERENT ECONOMIC TRADE-OFF.
Faster output does not automatically mean cheaper output, and capacity planning changes when a pipeline can finish work more quickly.
Data Studios calculation: if a workload consumed the same number of billable tokens per completed image or edit, the published token rates imply roughly 2× the spend for a fixed number of comparable jobs on Flare or Sunburst versus GPT-Image-2.
At the same time, a strictly serial and latency-limited pipeline that actually realizes a 50% latency reduction could theoretically complete about twice as many jobs in the same period, because cutting time per job in half doubles the maximum serial job rate.
Combining those two effects gives an upper-bound planning scenario: 2× price per token multiplied by as much as 2× serial throughput can raise the spend rate toward 4× if the pipeline remains continuously saturated and token use per job is unchanged.
That 4× figure is a derived scenario, not an OpenAI pricing claim: real spend depends on image dimensions and quality, token consumption, concurrency, retries, accepted-output rate, editing depth and whether the application is constrained by model latency in the first place.
The more useful operational metric is therefore not just cost per million tokens but cost per accepted production asset, paired with median generation time, retry rate and the number of editing turns required to reach the final result.
Images 2.5 ultimately makes the image stack more editor-like and potentially much faster, but its API pricing places an explicit premium on those gains; the strongest economic case will be in workflows where faster iteration and better edit preservation remove enough failed generations, manual rework or waiting time to justify the higher token rate.
··········
FOLLOW US FOR MORE.
DATA STUDIOS
datastudios.org

