GPT-Image 2.5 Release: Signal vs Noise
Lukas Vogel
Applied Research Editor

TLDRGPT-Image 2.5 launched on September 8, 2026, with faster generation, sharper image quality, more consistent edits, ChatGPT tools, and two API variants: Flare and Sunburst.
GPT-Image 2.5 Release Signals: Signal vs Noise
GPT-Image 2.5 launched on September 8, 2026. OpenAI released it in ChatGPT as ChatGPT Images 2.5 and in the API as two models: GPT-Image-2.5 Flare and GPT-Image-2.5 Sunburst.
TLDR GPT-Image 2.5 is now a confirmed OpenAI image-generation and image-editing release. It brings up to 50% lower generation latency than Images 2.0, sharper detail, more natural lighting and textures, stronger reference-subject preservation, and better consistency across repeated edits. ChatGPT also adds Sketch, templates, image comments, and prompt sharing. The remaining uncertainty concerns technical specifications, pricing details, standardized benchmarks, and how consistently the improvements generalize.
Updated 2026-09-04: GPT Image 2.5 was absent from the September 4 Astra announcement despite new desktop-build strings (see the Update below).
Key Takeaways
- OpenAI launched GPT-Image 2.5 on September 8, 2026, five days after the pre-launch signal documented in this article.
- ChatGPT Images 2.5 is rolling out across ChatGPT, ChatGPT Work, and Codex on desktop, mobile, and web.
- The API has two variants: GPT-Image-2.5 Flare, the default for most applications, and GPT-Image-2.5 Sunburst, which provides greater precision and control across edits with longer generation times.
- OpenAI’s launch materials describe up to 50% lower generation latency than Images 2.0, more natural lighting and textures, better preservation of recognizable subjects from reference photos, and stronger consistency across repeated edits.
- New ChatGPT tools include Sketch, templates, image comments, and prompt sharing.
- The pre-launch Arena names
mona-lisa-1andluna-lisa-alpharemain useful historical clues, but the supplied evidence does not establish them as the final public model names. - A standardized aggregate benchmark, full technical specification, and complete official pricing record are not included in the supplied evidence.
GPT-Image 2.5 at a Glance
| Specification | Current information |
|---|---|
| Status | Launched September 8, 2026 |
| ChatGPT availability | ChatGPT, ChatGPT Work, and Codex on desktop, mobile, and web |
| API availability | GPT-Image-2.5 Flare and GPT-Image-2.5 Sunburst |
| Default API model | GPT-Image-2.5 Flare |
| Precision API model | GPT-Image-2.5 Sunburst |
| Generation latency | Up to 50% lower than Images 2.0 |
| ChatGPT tools | Sketch, templates, image comments, and prompt sharing |
| Pricing | A launch-day report said API prices were unchanged from GPT-Image-2; complete official pricing details are not established in the supplied evidence |
| Benchmark status | Reported Arena placements, but no standardized numerical benchmark table in the supplied evidence |
| Context window | Not yet confirmed |
| Architecture and parameters | Not yet confirmed |
| License and open weights | Not yet confirmed |
What Was Actually Seen
The pre-launch signal bundle contained 6 topic posts from 3 authors, with 3 posts marked as evidence. That was enough to justify monitoring before release. It was not, at that stage, enough to establish a launch. The September 8 announcement and rollout resolved the central status question.
The first public thread in this bundle appeared on August 3. Mark Kretschmann asked where reports of an improved OpenAI image model came from. That post accumulated 19 replies, 9 reposts, 248 likes, and 14,236 views. Those numbers measure attention around the rumor, not model quality. They also show that the discussion predates the later release speculation by 31 days. The original August 3 question
The more concrete clue arrived on August 22. Kretschmann described two models “in the pipeline”: mona-lisa-1, presumed to be GPT-Image 2.5, and luna-lisa-alpha, presumed to be GPT-Image 2.5 Mini. He characterized the larger model as noticeably better than GPT Image 2, though not spectacularly so. He described the smaller model as faster and roughly comparable to the current model. He also said both still showed noise artifacts. That two-model report
This produced the first useful coined label: Two-Checkpoint Pipeline. Two-Checkpoint Pipeline means the community treated two anonymous Arena identifiers as separate candidates in one possible product family. It was a classification for the evidence, not an OpenAI product term. The launched product instead uses the public API names Flare and Sunburst.
On September 3, the release signal became more direct. TestingCatalog quoted language describing “higher-quality results, faster generation, and smarter creative tools,” and said a new image model had been seen in Image Arena testing. The post asked whether the model was GPT-Image 2.5. The September 3 upgrade report

Source: @testingcatalog
Four minutes later, Kretschmann wrote that a major image-model update was imminent for ChatGPT. He attributed improved quality, fewer noise artifacts, and higher speed to GPT-Image 2.5, while leaving open the possibility that Astra, or both models, could be released instead. That conditional wording mattered at the time. The post was a release forecast, not a release notice. The forecast was followed by the September 8 launch. The imminent-update claim
A separate post from Guizang described a new ChatGPT popup that appeared to signal an image-model upgrade. The screenshot did not identify a model ID or prove that GPT-Image 2.5 was already available, but the later rollout confirmed that the product-level signal preceded the launch. The popup observation
The pre-launch evidence pattern can therefore be summarized as Popup-Plus-Arena: a product-interface signal arriving alongside anonymous model testing. Popup-Plus-Arena was stronger than a single benchmark rumor, but weaker than an official model card or changelog entry. The launch converted that pattern into a confirmed product release.
What Shipped
GPT-Image 2.5 is available in ChatGPT as ChatGPT Images 2.5. The rollout covers ChatGPT, ChatGPT Work, and Codex on desktop, mobile, and web. The API launch added GPT-Image-2.5 Flare and GPT-Image-2.5 Sunburst.
Flare is the default for most applications. OpenAI’s developer guidance describes it as the starting point for general use, with the release’s quality, editing, and speed improvements. Sunburst is intended for cases that need greater precision and control across edits; its generation times are longer.
The launch materials describe up to 50% lower generation latency than Images 2.0. They also describe more natural lighting, richer textures, more recognizable subjects from reference photos, better detail preservation across repeated edits, more precise comment-based editing, and improved handling of complex layouts, styles, and transparent backgrounds.
ChatGPT adds several workflow features:
- Sketch: Users can draw directly in ChatGPT and use the drawing as a visual guide with
@Sketch. - Templates: Users can start from formats such as posters and merchandise designs.
- Image comments: Users can comment directly on an image to target a specific change.
- Prompt sharing: Prompts can be shared alongside images for others to remix.
The models were also reported live on OpenRouter, Replicate, and Higgsfield. Those integrations expand access beyond the ChatGPT and API surfaces, although the supplied evidence does not define identical limits, pricing, or feature support across each platform.
Why the Signal Matters
Image-model upgrades are difficult to evaluate through one attractive sample. The useful questions concern repeated behavior: whether typography survives at small sizes, whether lighting remains natural, whether edits preserve identity, and whether noise artifacts appear under difficult prompts.
The community reports focused on exactly those failure modes. A separate August account described mona-lisa-1 as producing less “AI plastic” appearance than GPT Image 2 High, including more natural overexposure, skin texture, lighting, fabric, and eye reflections. The same account claimed an image was detected by OpenAI’s SynthID tool, but that provenance claim came from one source and did not identify a specific model checkpoint. The qualitative comparison
This created a second label, Qualitative Realism Lift. Qualitative Realism Lift describes reported improvements in photographic texture and lighting without a shared scoring protocol. After launch, those reported directions align with the stated improvements in lighting, textures, and reference-subject preservation, but the label remains a description of observed behavior rather than a benchmark result.
The artifact question is more important than the upgrade language. One August report said luna-lisa-alpha looked more realistic and that earlier grain appeared to be gone. Another report said the alleged model failed a “centaur doing a handstand” prompt against GPT Image 2, Nano Banana Pro, and Grok Imagine Image 2.0. Kretschmann later said both alleged new models still suffered from noise artifacts. The reported centaur test
The post-launch evidence remains mixed. Some reports describe broken or smeared-looking results as improved, while other tests still found visible artifacts. These outcomes are not necessarily contradictory: noise can vary by prompt, rendering mode, checkpoint, or sampling path. The reports do not disclose enough shared settings to isolate the cause. The right label remains Noise Artifact Question: an important quality dimension that should be tested directly rather than described as fully solved.
On the evidence available after launch, GPT-Image 2.5 targets realism, speed, and artifact reduction, but the noise claim is an improvement direction rather than a guarantee of artifact-free output.
How to Read the GPT-Image 2.5 Benchmark Claims
A public launch does not automatically create a standardized benchmark. The supplied evidence still does not provide a reproducible GPT-Image 2.5 benchmark table with shared prompts, scoring rules, sample coverage, and variance estimates.
One community post shared a reproducible-looking comparison configuration: GPT-Image 2.5 at medium quality, 1K resolution, with no thinking, compared with GPT Image 2, Grok Imagine Image 2, and Reve 2.1. The post did not publish numerical results. The structured comparison setup
This is a primary test report, but it is not a benchmark table. The difference is material. A test configuration tells readers how an observation was generated. A benchmark requires published outputs, scoring rules, sample coverage, and enough repetition to estimate variance.
The launch also generated Arena-related claims. One commentary account reported Sunburst at number one and Flare at number two across the Text-to-Image, Image Edit, and Multi-Image Edit Arenas. The supplied evidence does not include the underlying score table or a standardized, independently reproducible evaluation, so those placements should be treated as reported Arena results rather than a complete benchmark record.
GPT Image 2 provides a limited baseline. A third-party evidence tracker reported that it sat first on a public text-to-image leaderboard at 1,381±5 Elo, across 76 listed models. The same tracker said GPT Image 2 was listed as “medium” and that it could return up to 8 images per prompt, with an output ceiling of up to 2K. Those details describe the baseline, not GPT-Image 2.5.
The comparison gap is now narrower but remains measurable only in part:
| Measure | GPT-Image 2.5 signal | GPT Image 2 baseline |
|---|---|---|
| Public model identity | Launched as ChatGPT Images 2.5, GPT-Image-2.5 Flare, and GPT-Image-2.5 Sunburst | Identified in the supplied evidence |
| Aggregate score | Reported Arena placements, but no standardized numerical table supplied | Reported at 1,381±5 Elo |
| Generation speed | Up to 50% lower latency than Images 2.0 is stated in launch materials | No comparable speed number in this signal set |
| Test resolution | 1K in one community setup | Up to 2K output ceiling reported |
| Batch output | Not confirmed in the supplied evidence | Up to 8 images per prompt reported |
| Noise behavior | Improved in some tests, with mixed results and residual artifacts | Described as a current-model issue |
The phrase Benchmark Vacuum still captures this state. Benchmark Vacuum means a launched model can attract detailed qualitative discussion while lacking a public, matched, reproducible evaluation set.
A precise claim such as “GPT-Image 2.5 scores X on Arena” would therefore remain unsupported unless accompanied by a traceable score and methodology. The defensible comparison is more specific than before: launch materials state a latency improvement and several workflow improvements; community reports describe stronger editing consistency; the evidence does not establish how large, consistent, or general the gains are across all tasks.
GPT-Image 2.5 vs GPT Image 2: What the Signal Says
GPT Image 2 is the most relevant competitor because the pre-launch reports and post-launch commentary use it as the direct baseline. GPT Image 2.5 is now a public release with ChatGPT and API availability, while the comparison evidence remains uneven across quality, speed, and benchmark dimensions.
Identity and availability. GPT Image 2 is named in the supplied evidence and is described as real. GPT-Image 2.5 launched on September 8, 2026, as ChatGPT Images 2.5 and as the Flare and Sunburst API variants. The earlier anonymous codenames were pre-launch clues, not the confirmed public product names.
Image quality. Launch materials describe sharper detail, more natural lighting and textures, and better preservation of recognizable subjects from reference photos. Community testers also described cleaner text and stronger visual control. These observations are directionally consistent, but they are not matched scores across a published test set.
Speed. The launch materials state up to 50% lower generation latency than Images 2.0. Flare is the faster default-oriented API model, while Sunburst is designed for more precise work and has longer generation times. The evidence does not provide hardware conditions, image-size settings, percentile measurements, or a complete latency distribution.
Editing consistency. The clearest post-launch improvement is control across edits. The release emphasizes detail preservation across repeated edits and comment-based local changes. Community reports also described stronger identity, texture, and background retention through successive edits. Those results support the workflow claim, but do not imply deterministic output.
Noise artifacts. Fewer artifacts are part of the GPT-Image 2.5 improvement story. Yet the post-launch evidence remains mixed, with some reports describing improvement and other tests still finding checkerboard, broken, or smeared-looking results. GPT Image 2 remains the baseline for this issue, but the bundle supplies no standardized artifact rate for either model.
Evaluation evidence. GPT Image 2 has a reported 1,381±5 Elo score. GPT-Image 2.5 has reported Arena placements but no aggregate numerical score table in the supplied evidence. At the reported 1K, medium-quality test setting, GPT-Image 2.5 is not shown to beat GPT Image 2 because the comparison post publishes no numerical result.
What We Know vs. What We Don't
The cleanest reading separates confirmed launch facts from pre-launch observations and unresolved technical claims.
What the evidence documents
- OpenAI launched GPT-Image 2.5 on September 8, 2026.
- ChatGPT Images 2.5 is rolling out across ChatGPT, ChatGPT Work, and Codex on desktop, mobile, and web.
- The API includes GPT-Image-2.5 Flare and GPT-Image-2.5 Sunburst. Flare is the default for most applications; Sunburst is intended for greater precision and control across edits and takes longer to generate.
- Launch materials describe up to 50% lower generation latency than Images 2.0, more natural lighting and textures, stronger reference-subject preservation, better repeated-edit consistency, comment-based editing, and improved handling of complex layouts, styles, and transparent backgrounds.
- ChatGPT adds Sketch, templates, image comments, and prompt sharing.
- The pre-launch Arena names
mona-lisa-1andluna-lisa-alphawere reported by community observers. The supplied evidence does not establish either as the final public model name. - GPT-Image 2.5 was reported live on OpenRouter, Replicate, and Higgsfield.
- A launch-day report said API prices were unchanged from GPT-Image-2, but the supplied evidence does not include a direct official pricing table or complete confirmed token rates.
- One commentary account reported Sunburst at number one and Flare at number two across several Arena categories, but the supplied evidence does not provide a standardized score table.
What We Don't Know Yet
- What are the architecture, parameter count, and training details? The supplied evidence does not establish the model internals, parameter count, training method, or training-data figure.
- What is the context window and full API specification? The launch confirms the two API variants, but the supplied evidence does not establish context limits, rate limits, input constraints, or all supported settings.
- What are the complete official prices? A launch-day report said API prices were unchanged from GPT-Image-2, but complete official pricing details are not included in the supplied evidence.
- How deterministic is the model? Consistency tests report substantial improvement but still describe residual jitter and nondeterministic behavior.
- How broadly do the improvements generalize? The evidence is strongest for repeated editing, reference-subject preservation, text rendering, localized changes, and selected creative workflows. It does not establish performance across every prompt type.
- Has the noise artifact problem been solved? The available reports indicate improvement in some cases, but artifacts remain in other tests and no standardized rate is supplied.
- What license terms and weight-access terms apply? The supplied evidence does not establish a license or indicate that open weights were released.
FAQ
Is GPT-Image 2.5 released?
Yes. OpenAI launched GPT-Image 2.5 on September 8, 2026. It is available in ChatGPT as ChatGPT Images 2.5 across ChatGPT, ChatGPT Work, and Codex on desktop, mobile, and web, and in the API as GPT-Image-2.5 Flare and GPT-Image-2.5 Sunburst.
What are the GPT-Image 2.5 API models?
The two API models are GPT-Image-2.5 Flare and GPT-Image-2.5 Sunburst. Flare is the default for most applications and combines the quality, editing, and speed improvements. Sunburst is designed for greater precision and control across edits, with longer generation times.
What improved in GPT-Image 2.5?
The launch materials describe up to 50% lower generation latency than Images 2.0, more natural lighting and textures, better preservation of recognizable subjects from reference photos, stronger consistency across repeated edits, and more precise comment-based editing. The release also adds Sketch, templates, and prompt sharing in ChatGPT.
Where can I use GPT-Image 2.5?
ChatGPT Images 2.5 is rolling out across ChatGPT, ChatGPT Work, and Codex on desktop, mobile, and web. The API variants are available through the API, and GPT-Image 2.5 has also been reported live on OpenRouter, Replicate, and Higgsfield.
What happened to the Arena codenames?
The pre-launch codenames mona-lisa-1 and luna-lisa-alpha were reported as anonymous Arena checkpoints. They were useful release clues, but the supplied evidence does not establish them as the final public model names. The launched API names are GPT-Image-2.5 Flare and GPT-Image-2.5 Sunburst.
Has GPT-Image 2.5 fixed noise artifacts?
Not completely. Launch-day reports describe improvement in broken or smeared-looking results, but testing still found mixed results and persistent artifacts in some cases. The supplied evidence does not establish a standardized artifact rate.
Does GPT-Image 2.5 have a benchmark score?
The supplied evidence does not provide a standardized published numerical benchmark for GPT-Image 2.5. One commentary account reported Sunburst at number one and Flare at number two across several Arena categories, but the bundle does not include a reproducible score table.
What does GPT-Image 2.5 cost?
A launch-day report said the API prices were unchanged from GPT-Image-2, but the supplied evidence does not include a direct official pricing table or complete confirmed token rates. Treat detailed pricing claims as unconfirmed until checked against OpenAI's current API terms.
What is still unknown about GPT-Image 2.5?
The supplied evidence does not establish the model architecture, parameter count, context limits, training-data details, license terms, deterministic behavior, or a standardized evaluation across the full range of image-generation and editing tasks.
What This Means for Builders
Builders can now evaluate GPT-Image 2.5 directly rather than preparing for a possible release. The available evidence supports a structured comparison, not an automatic migration decision.
First, freeze a prompt set that targets the reported improvement areas. Include small typography, multilingual text, skin and fabric texture, hard directional light, repeated edits, and identity preservation. The supplied evidence specifically discusses realism, overexposure, eye reflections, fabric, and noise. Those categories are more useful than a gallery of unrelated hero images.
Second, record generation settings. The only structured comparison in the bundle used medium quality, 1K output, and no thinking. A comparison should preserve those settings for the baseline, then add a second pass at another available quality or resolution. Without that control, a speed or quality claim can reflect configuration rather than model behavior.
Third, separate output quality from throughput. “Faster generation” is supported by the launch materials, including the up-to-50% lower latency claim, but a production test should still capture time to first result, total time for all requested images, failure rate, and output size. GPT Image 2 is reported to return up to 8 images per prompt and up to 2K output, so batch behavior should be logged separately from single-image latency.
For a controlled baseline, builders can also test GPT Image 2 against the same prompt suite and compare it with both Flare and Sunburst.
The fourth step is provenance and artifact inspection. The single SynthID-related claim should not be treated as model identification. Save original files, metadata, and prompt settings. Test repeated generations rather than one image. A visual artifact that disappears in one sample may return under another composition, lighting condition, or edit chain.
The fifth step is to separate confirmed availability from unconfirmed economics. GPT-Image 2.5 is available in ChatGPT and through the API, with additional live reports for OpenRouter, Replicate, and Higgsfield. However, the supplied evidence does not provide a complete official pricing table, rate-limit schedule, or identical feature matrix across those services. Teams should verify those terms before building a production cost model.
The Week Ahead
The next useful evidence will be less dramatic than another screenshot. A complete API specification, official pricing table, model card, or reproducible benchmark would resolve the remaining technical questions. A public Arena score with methodology would create a more measurable comparison. A larger image set would clarify whether the reported realism and editing gains survive beyond individual examples.
Three signals deserve priority:
- Watch the model documentation. Check whether OpenAI publishes context limits, supported settings, rate limits, pricing, and technical details for Flare and Sunburst.
- Run the same 1K, medium-quality prompt suite. Compare typography, texture, edits, speed, and noise against GPT Image 2 without changing settings.
- Pin the artifact definition. Record whether “less noise” means fewer visible pixels, fewer failed generations, or simply cleaner samples in selected prompts.
GPT-Image 2.5 now has a verified launch and a clear product surface. The remaining work is evaluation: determining how large, consistent, and general the stated improvements are, and how they translate into production cost and reliability.
Update — 2026-09-04
As of September 4, GPT Image 2.5 had still not been confirmed in OpenAI’s public Astra announcement. That absence weakens the September 3 same-day launch speculation, but does not rule out a later release. The September 4 signal review
Separate third-party reports identified GPT-Image-2.5-related strings in a ChatGPT/Codex Desktop build dated September 1, while OpenAI’s public changelog remained unchanged. This is additional product-build evidence, not an official model announcement or proof that the model is available. The build-string report
Building similar image-generation workflows? On kie.ai you can try GPT Image 2.5, Qwen Image 2.1, and Grok Imagine Image 2.0.
About Lukas Vogel
Lukas reads the papers and model cards so you do not have to, focusing on reproducible claims.
View all posts by Lukas Vogel