GPT Image 2.5 vs GPT Image 2

Same prompts, same rubrics, same defaults. Only the model changed.

OpenAI released GPT Image 2.5 on 8 September 2026 as two API tiers: Flare, the fast default, and Sunburst, the precision tier. The launch notes promise sharper detail, more precise multi-step edits and stronger multi-turn consistency. We already hold graded, published GPT Image 2 results on fixed prompts from July, so the cheapest honest question is the one a buyer actually has: on the tests the old model failed, what changed? Both tiers ran three times per test the day after launch, at defaults, no retries. Every run is below with its mark.

Sunburst learned mirrors. Neither tier learned refraction. Flare is a small step backwards.

Sunburst is the first OpenAI image model to pass an implicit-physics test here: two of three two-mirror runs reflect the person’s front, where GPT Image 2 and Flare show the back every time. Nothing in any version reverses the arrow seen through water. On prompt adherence the ceiling is unchanged — Sunburst ties GPT Image 2 on the fruit bowl (57.1) and the duck ladder (100) — while Flare scores below the April model on both (52.4 and 95.8) and drops a character’s jacket colour in one consistency run.

Scoreboard

Per-run marks, then a total /100 per test. The GPT Image 2 column is its original published result, locked before 2.5 existed, not a re-run. The four-panel row is observed, not rubric-scored.

TestGPT Image 2 (April 2026)GPT Image 2.5 FlareGPT Image 2.5 Sunburst
Two-mirror corner (implicit physics)

PASS = each mirror shows the person's front. The prompt never says what the mirrors should show.

0

FAIL · FAIL · FAIL

0

FAIL · FAIL · FAIL

67

PASS · PASS · FAIL

Arrow refraction (implicit physics)

PASS = the arrow seen through the glass of water is reversed. The prompt never mentions refraction.

0

FAIL · FAIL · FAIL

0

FAIL · FAIL · FAIL

0

FAIL · FAIL · FAIL

Wrong-coloured fruit bowl (prompt adherence)

7 marks per run: banana pink · lemon · lime · orange skin · orange flesh · strawberry colour · strawberry rot.

57.1

4/7

52.4

3/7 · 4/7 · 4/7

57.1

4/7 · 4/7 · 4/7

Duck transform ladder (8 edits)

One mark per edit: ✓ done as asked · ½ right edit, flawed · ✗ missing, out of order, or changed in a way not asked for.

100

8/8 · 8/8 · 8/8

95.8

7/8 · 8/8 · 8/8

100

8/8 · 8/8 · 8/8

Same courier, four panels (consistency, not scored)

Does the invented courier stay the same person across all four panels? Observed, not rubric-scored.

1/1

consistent

2/3

jacket changes colour in panel 3 · consistent · consistent

3/3

consistent · consistent · consistent

Two-mirror corner (implicit physics)

Sunburst is the first OpenAI image model to pass this test — two runs out of three reflect the person's front. Flare, like GPT Image 2, puts the person's back in the mirror every time.

Two-mirror corner (implicit physics): every run for GPT Image 2, GPT Image 2.5 Flare and GPT Image 2.5 Sunburst, each with its mark.

PASS = each mirror shows the person's front. The prompt never says what the mirrors should show. Full benchmark, other models and the verbatim prompt →

Arrow refraction (implicit physics)

Nothing moved. All nine runs render a perfect scene and leave the submerged arrow pointing the same way as the one above it.

Arrow refraction (implicit physics): every run for GPT Image 2, GPT Image 2.5 Flare and GPT Image 2.5 Sunburst, each with its mark.

PASS = the arrow seen through the glass of water is reversed. The prompt never mentions refraction. Full benchmark, other models and the verbatim prompt →

Wrong-coloured fruit bowl (prompt adherence)

The ceiling is unchanged. Sunburst is the first GPT model to recolour the lemon, but it then leaves the orange's skin natural; no run from any version has ever recoloured the lime or the strawberry. Flare's 52.4 is a quality-tier effect: re-run at quality high (which bills exactly what GPT Image 2's default bills) it scores 4, 4, 4 — the old model's 57.1 to the decimal.

Wrong-coloured fruit bowl (prompt adherence): every run for GPT Image 2, GPT Image 2.5 Flare and GPT Image 2.5 Sunburst, each with its mark.

7 marks per run: banana pink · lemon · lime · orange skin · orange flesh · strawberry colour · strawberry rot. Full benchmark, other models and the verbatim prompt →

Duck transform ladder (8 edits)

Sunburst keeps the clean sweep, mirror reflection included. Flare drops one mark: on run 1 the face-away duck came back yellow, a change nobody asked for.

Duck transform ladder (8 edits): every run for GPT Image 2, GPT Image 2.5 Flare and GPT Image 2.5 Sunburst, each with its mark.

One mark per edit: ✓ done as asked · ½ right edit, flawed · ✗ missing, out of order, or changed in a way not asked for. Full benchmark, other models and the verbatim prompt →

Same courier, four panels (consistency, not scored)

OpenAI's headline claim is better multi-turn consistency. On the one consistency test we have, Sunburst holds the character in all three runs; Flare loses the courier's red jacket in one panel of one run.

Same courier, four panels (consistency, not scored): every run for GPT Image 2, GPT Image 2.5 Flare and GPT Image 2.5 Sunburst, each with its mark.

Does the invented courier stay the same person across all four panels? Observed, not rubric-scored. Full benchmark, other models and the verbatim prompt →

Flare vs Sunburst: speed and price

Measured through the same gateway, in the same hour, on the same prompts, with three fresh GPT Image 2 generations as a control. Text-to-image at 1536×1024: GPT Image 2 averaged 30.9s, Flare 12.5s (median 12s), Sunburst 14.6s (median 14.1s). The eight-duck edit: Flare 12.3s, Sunburst 16.9s. So OpenAI’s “up to 50% lower latency” claim for Flare holds here (about 60% lower than GPT Image 2), and Sunburst costs you roughly a fifth more waiting than Flare, not a return to the old speed.

Price: the gateway billed both 2.5 tiers the same at default quality, about $0.014 per generation and $0.03 per edit, against about $0.055 for a GPT Image 2 generation at its default. That gap is a quality-tier difference (2.5 defaults to “medium”; the older model’s default bills like a higher tier), which means the scores above compare the new tiers at a cheaper setting than the April baseline. To check that this doesn’t manufacture a regression, Flare was re-run on the fruit bowl at quality “high” (billed $0.054, the same as GPT Image 2’s default): 4/7, 4/7, 4/7 — exactly the old model’s 57.1. So at matched spend Flare is the April model; at default spend it is a shade below it.

How it’s scored

  • Each test keeps its original, published rubric: binary PASS/FAIL per run for the two implicit-physics tests; seven marks per run for the fruit bowl; one mark per edit for the duck ladder. The four-panel consistency test has no rubric and is reported as an observation only.
  • Three runs per tier per test, identical prompts, API defaults (quality medium, 1536×1024), no retries, no cherry-picking. Every run is on the composites above with its mark; no number comes from a run you can’t see.
  • The GPT Image 2 column reuses the published July results rather than re-running the model, so the baseline was fixed before 2.5 existed. Its fruit-bowl entry is that benchmark’s single published run, disclosed as such.
  • Access, per model: Published July results · API defaults · 1536×1024 · AIMLAPI openai/gpt-image-2.5-flare · defaults (quality medium) · 1536×1024 · 2026-09-09 · AIMLAPI openai/gpt-image-2.5-sunburst · defaults (quality medium) · 1536×1024 · 2026-09-09.
  • Grading is by hand against the written rubric, fine detail checked zoomed in per slot. The two Sunburst mirror passes are the finding that matters most, so they are published at full size on the composite for re-grading.
  • Free to reuse with attribution (CC BY 4.0). Press & data requests: promptfrenzy.com/press.

Run these prompts yourself

Every test above links to its benchmark page with the verbatim prompt and a free run-it-yourself button across the models we host. No setup.

All benchmarks →