GPT Image 2.5: the image flagship that writes multi-round editing consistency into the API contract
gpt-image-2-5
OpenAI's image generation and editing model released 2026-09-08, in two variants: Flare (the default, higher quality than GPT Image 2 at 50% lower latency) and Sunburst (the precision tier, trading longer generation time for fidelity on intricate detail). Four endpoints in total, each variant with text-to-image and edit, identical API surface and identical pricing. Capability surface: 3840px long edge, six quality levels (auto/low/medium/high/xhigh/max, two above the previous ceiling), editing with up to 16 reference images per request, optional mask, and real alpha transparency on PNG/WebP. The launch weight sits on consistent details across multiple edits and comment-based edits that change only what you ask, i.e. multi-round editing that does not drift. Third-party reading: on the Artificial Analysis arena this site syncs (2026-09-22), sunburst is #1 on both text-to-image (Elo 1423) and image editing (1526), with flare #2 on both (1401/1482), the only family holding two crowns at once; both rows are still preliminary, however, with vote counts (8,791/8,014) far below the gpt-image-2 row they overtake (82,862). Closed, reachable only through ChatGPT and the API, no self-hosting or fine-tuning. Not benchmarked by us; graded C (vendor-stated).
- CONFIDENCE
- Vendor Claim
- Official model card or keynote only, no independent re-test
- KEY METRIC
- Arena 文生图 / 图像编辑 Elo 双榜第 1
- Vendor Claim · 2026-09
- MATURITY
- Product
- research → demo → product → production
Our takeWe grade this C (vendor-stated), but the reason it holds the top of the image ladder needs stating precisely: it is not there because it carries the OpenAI name, it is there because it holds the #1 rank on two third-party arenas at once and carries the newest release date. Both are checkable, neither is an adjective.
Three things on the capability surface carry real production weight. The first is multi-round editing that does not drift: the demonstrated sequence is three consecutive edits, each demanding that nothing else change, and the pass criterion is that the previous round's edit survives into the next. For an agent workflow this is the decisive link, because once an editing tool is called repeatedly by an agent, cumulative drift decouples the fifth round's output from the first round's intent. This single property decides whether the tool can be handed to an agent to run unattended. The second is 16 reference images per edit, which turns "compose several subjects into one frame" from a multi-pass img2img chain into one call. The third is quality expanding from three levels to six (xhigh and max above the old ceiling): together with the Flare/Sunburst split, "fast" and "detailed" become two choices on one price sheet rather than a model-selection decision.
Two boundaries belong on the front of the card. The first is that the crowns rest on unsettled vote counts: both the sunburst and flare rows were still preliminary at scrape time, with 8,791 and 8,014 text-to-image votes against 82,862 for the gpt-image-2 (medium) row they overtake and 157,906 for nano-banana-pro, at a ±9 confidence interval. The accurate claim is therefore "#1 on both boards on early votes", our board rows refresh with the source, and this assessment must be re-checked when they do. The second is that it is closed, and data leaves your perimeter: no self-hosting, no fine-tuning. For private-deployment teams the image candidate set remains the open-weight tier, Qwen-Image-2.1 / FLUX / Wan, not this. The two routes are not competitors; they are the optimum under two different constraints.
The problem it addresses: making multi-round editing a contract, not a hope
By 2026 image generation no longer lacks "a good-looking image". What actually blocks a production line is continuity: when you issue the second edit, does the first edit survive it; after a style change, is the person in the reference photo still the same person; on a transit map crowded with small type, can you still read the station names once you crop in. GPT-Image-2.5 puts nearly its whole launch weight on those three questions. OpenAI's announcement of 2026-09-08 lists four improvements: faster generation, improved fidelity for more natural and recognisable images, consistent details across multiple edits, and comment-based edits that change only what you ask for.
Splitting one generation into two variants is the most practically useful structural decision here. Flare is the default tier: OpenAI's framing is higher image quality than GPT Image 2 at 50% lower latency, aimed at high-volume and interactive work. Sunburst is the precision tier: same surface, text-to-image and editing of existing images, but it spends longer per request in exchange for extra fidelity on intricate detail, aimed at images that will be viewed large or cropped into. The API surface is identical and the pricing is identical, so choosing between them is a latency trade, never a cost trade.
Capability surface: the numbers you can actually schedule against
| Capability | Reading | Production meaning |
|---|---|---|
| Max resolution | 3840px long edge; both edges multiples of 16, aspect ratio up to 3:1, total pixels 655,360 to 8,294,400 | Native 4K (3840x2160) with no upscaling pass; but the pixel floor means tiny outputs are out of scope |
| Quality levels | auto / low / medium / high (default) / xhigh / max | Two levels above the GPT Image 2 ceiling; xhigh and max are where Sunburst earns its latency |
| Edit references | Up to 16 reference images per edit; up to 4 images per request; optional mask | Tasks like "compose six portraits into one group photo" land in one pass; the mask is optional, not required |
| Transparent background | Real alpha channel on PNG / WebP output | Product shots and icons drop straight onto a layout, removing a rotoscoping step; JPEG carries no alpha, so do not pair it with transparency |
| Endpoints | Four: Flare and Sunburst, each with text-to-image and edit | Same parameters, same price sheet; migration cost is the model name |
"Edits that carry across turns" deserves its own paragraph because it is a testable engineering behaviour rather than an adjective. The demonstrated sequence is: round one reupholsters a sofa in deep forest-green velvet with the instruction "change nothing else in the room"; round two hangs a framed abstract landscape on the bare wall above it with "leave the rest of the room exactly as it is"; round three moves the time of day to evening while demanding the furniture, the artwork and the layout stay identical. The pass criterion across three rounds is that the previous round's change is still there. That is precisely the hardest link in an agent workflow: once an editing tool is called repeatedly by an agent, cumulative drift decouples the fifth round's output from the first round's intent.
Third-party reading: number one on two boards, with the caveats stated
On the Artificial Analysis arena this site syncs (scraped 2026-09-22), GPT-Image-2.5 is the only model family holding the top rank on both text-to-image and image editing:
- Text-to-image (arena-t2i): sunburst #1, Elo 1423; flare #2 at 1401; the previous generation gpt-image-2 (medium) #3 at 1381.
- Image editing (arena-image-edit): sunburst #1, Elo 1526; flare #2 at 1482; gpt-image-2 (medium) #3 at 1461.
- What follows is not the same vendor: on text-to-image, #4 is Microsoft's mai-image-2.6 (1334), #5 xAI's grok-imagine-image-2.0 (1302), #7 Meta's muse-image (1276), #9 Google's nano-banana-2 (1260), #10 ByteDance's seedream-5.0-pro (1256), #11 Alibaba's qwen-image-3.0-pro (1254); on image editing, #4 is grok-imagine-image-2.0 (1430) and #5 mai-image-2.6 (1429).
The honest half: at scrape time both the sunburst and flare rows still carried preliminary=true. Their text-to-image vote counts were only 8,791 and 8,014, against 82,862 for the gpt-image-2 (medium) row they overtake. The editing board is better (29,252 and 27,048 votes) but still short of gpt-image-2's 259,008 and nano-banana-pro's 582,162. This is therefore a number one on early vote counts, with a ±9 confidence interval, not a settled verdict. We place it at the top of the image SOTA ladder because it holds both crowns at once and carries the newest official release date (2026-09-08), not because its vote totals have matured.
Cost: quality level is the only real lever
fal measured actual billing across the full size-by-quality matrix on 2026-09-08 with short prompts, and Flare and Sunburst price identically. Text-to-image at the default high quality: about $0.0362 for 1024x768, $0.0528 for 1024x1024, $0.0395 for 1920x1080, $0.1002 for 3840x2160 (4K, the maximum). Medium runs $0.0091 to $0.0260, low starts at $0.0041. Editing the same sizes costs slightly more because the reference image bills as image input: $0.0445 to $0.1084 at high, with a single 1024x1024 reference adding about $0.008 and a 4K reference about $0.012, and more references stacking on top. Underneath, the model bills by token: text input $5.00 per million, image input $8.00 per million, image output $30.00 per million, cached input $1.25 and $2.00, total rounded up to the nearest $0.0001. The scheduling conclusion: fix size first, then quality, because quality sets the order of magnitude (high costs roughly four times medium) and xhigh/max go further still.
Boundaries
- Closed, API and ChatGPT only: reachable through ChatGPT (including ChatGPT Work and Codex users) and the API. No self-hosting, no fine-tuning, data leaves your perimeter. For private-deployment teams the candidate set remains the open-weight tier: Qwen-Image-2.1, FLUX, Wan.
- Sunburst latency is an explicit cost: OpenAI publishes no second count, only "longer generation times". Interactive work should default to Flare.
- Type still needs eyes on it: dense layouts (transit maps, dashboards, packaging labels) are the showcased strength, but mixed Chinese-Latin signage and small package copy still warrant a human pass in production.
- Votes have not settled: both crowns are preliminary rows and may move on a re-scrape in a few weeks. Our board rows refresh with the source, and this assessment must be re-checked when they do.
- 655,360 pixel floor: thumbnails and favicon-scale assets are not its target scene.
Our verification status
Facts on this page come from three read-in-full sources: OpenAI's developer community announcement Introducing GPT Images 2.5 in the API and ChatGPT (posted 2026-09-08T19:12:35Z, carrying the Flare/Sunburst split and the four improvements), fal's model page (four endpoints, resolution and quality-level constraints, transparent background, 16 reference images, the full measured price matrix), and the arena-t2i / arena-image-edit board rows already stored on this site (including vote counts and preliminary flags). We did not reproduce any benchmark: no side-by-side run over a fixed prompt set, no quantified drift rate across editing rounds, no measured latency gap between Flare and Sunburst. Model-capability claims are therefore graded vendor-stated (C), while the arena positions are third-party readings (between A and B depending on arena protocol and how far votes have settled). To lift this to grade A we would need: same-prompt Flare versus Sunburst comparisons on a fixed set, a structural-similarity measurement after three editing rounds, a matched generation batch against mai-image-2.6 / seedream-5.0-pro / nano-banana-pro, and one board re-check after votes mature.