PromptFrenzy Benchmark Report

The State of AI Image Models, September 2026

We run the same prompts through every major AI image model and publish the results side by side. For most of this year OpenAI was the slowest name in the field. With GPT Image 2.5, that stopped being true.

Data as of 9–10 September 2026.

OpenAI just closed the speed gap it had been losing on

GPT Image 2.5 Flare produces an image in 12.4s median, against 35.9s for GPT Image 2 — a 2.9× speed-up in a single version. That puts it level with Google's Nano Banana 2 at 12.3s, a difference well inside the run-to-run noise on both sides.

When we measured in June, OpenAI's image model was several times slower than Google's, and that gap was the most striking number in the report. It is gone. After a year of Google owning the speed argument, the two are now indistinguishable on it.

Time-to-image, by model

Four fixed text-to-image scenes, three runs per model per scene, one generation per request.

ModelMedianSamples
Nano Banana 2 (Gemini)12.3s16
GPT Image 2.5 Flare12.4s12
GPT Image 2.5 Sunburst14.4s12
ChatGPT (GPT Image 2)35.9s21

Method: wall-clock time from request to returned image, identical scenes, one shot per run, no retries counted. GPT Image 2.5 was measured on 9 September 2026; GPT Image 2 and Nano Banana 2 on the same scenes through the same harness on 10 September. A fifth image-editing scene was run but is excluded here, because Nano Banana 2 is not served on the image-edit endpoint and it would not have been like-for-like. Every row has at least 12 runs behind it.

What that looks like for a real user

Benchmarks run one clean request at a time. Real users hit queues, retries and busy hours. So we also measured every image generation on PromptFrenzy over the last 60 days — all prompts, real traffic:

ModelMedian waitGens
Nano Banana 2 (Gemini)12.7s610
ChatGPT (GPT Image 2)40.2s372

The medians track the benchmark closely, which is a good sign the lab numbers aren't flattering anyone. But the means sit far above the medians on both models: a minority of slow generations drags the average up, and on GPT Image 2 the slowest tenth take 80 seconds or more. The typical wait is better than the average number suggests.

GPT Image 2.5 is absent here because it is not yet serving production traffic on our platform, so we have no production sample for it. Internal, admin and automated-test accounts are excluded throughout.

The field is moving faster than the models render

Speed matters more this year because the roster keeps growing. Of the AI image and video models we track with a known launch date, a third shipped in the first half of 2026 alone — Google, OpenAI, xAI, Black Forest Labs and ByteDance all pushed new models in the last six months.

See the full release timeline →

Same prompt, every model

The most useful thing we publish isn't a number — it's the side-by-side. We take one brief and run it through every model, one shot each, no cherry-picking. The four scenes behind the table above are deliberately awkward: a refraction test, a reflection test, a compound-instruction test, and an open-ended four-panel comic. They are built to separate models rather than flatter them, and the differences in interpretation, text rendering and physical plausibility rarely match each model's reputation.

See the head-to-heads →

What it means

  • →The speed argument is over, for now. OpenAI and Google are level on time-to-image. Anyone still choosing between them on speed is working from last quarter’s data.
  • →Version jumps matter more than vendor reputation. GPT Image 2 and GPT Image 2.5 are further apart on speed than GPT Image 2.5 and its closest competitor.
  • →Averages hide the wait. Median and 90th-percentile latency tell a user far more than a mean, and almost nobody publishes them.
  • →There is still no single “best.” The right model depends on the task — which is the whole reason to compare on your own prompt rather than trust a leaderboard.

Methodology & how to cite

The benchmark table uses four fixed text-to-image scenes, three runs per model per scene, one generation per request, wall-clock from request to returned image. The production table covers every non-video generation on PromptFrenzy in the 60 days to 10 September 2026, across all prompts rather than fixed scenes — a measure of what users wait, not a controlled model comparison, which is why we report it separately. Speed reflects models as served to us, not vendor-direct lab numbers. Sample sizes are shown for every row; we do not publish a figure with fewer than 12 runs behind it.

Please cite as Source: PromptFrenzy — promptfrenzy.com. For a custom comparison, the underlying data, or a quote, contact maria@promptfrenzy.com. More for journalists on our press & data page.