Benchmark results

A comparison of twelve image generation models by Elo score from human preferences in August 2026 reveals the order of visual appeal. The data comes from a benchmark where evaluators compared image pairs without knowing the model that produced them. The table includes models from OpenAI, Reve, Microsoft AI, Google, HiDream, Bytedance, NVIDIA and Recraft. The measurement reflects subjective taste, not technical accuracy in following prompts, which is decisive for commercial selection.

At the top of the ranking stands GPT Image 2 with an Elo of 1 339, narrowly ahead of Reve 2.1 at 1 299. Third place goes to MAI-Image-2.5 with 1 270. The middle of the field comprises three models with very close scores: Nano Banana 2 Lite and GPT Image 1.5 both score 1 263, followed by Nano Banana 2 at 1 262. HiDream-O1-Image-1.5 reaches 1 245, Seedream 5.0 Pro 1 240. The bottom third opens with Nano Banana Pro at 1 225, Cosmos3-Super-Text2Image at 1 219, MAI-Image-2.5-Flash at 1 208 and Recraft V4.1 Utility at 1 205. Surprisingly, Google’s Lite and base Nano Banana 2 variants outperform the Pro version.

What Elo scores do not tell you

Elo expresses the probability that people prefer the output of one model over another, not the degree of adherence to a text prompt. A model with a higher Elo may generate aesthetically more pleasing images that nevertheless ignore key details of the prompt. Differences of tens of points, for example between 1 339 and 1 299, signal a clear preference in head-to-head comparison but say nothing about consistency across different styles or about adherence to specific compositions. For commercial use it is risky to rely on this ranking alone without internal tests on the specific prompt types required, because the benchmark set may not cover specialised domains such as technical illustration or brand graphics.

Choosing a model for production

For organisations generally, models at the top of the ranking such as GPT Image 2 or Reve 2.1 offer the greatest chance of visually attractive results for marketing and presentation purposes. If low latency or cost is the priority, models labelled Flash, Lite or Utility in the lower half of the table may represent a reasonable compromise. The choice should reflect whether the task is generating hero images where first impression decides, or serial production where precision and repeatability dominate.

Frequently asked questions

Which image generation model has the highest Elo score in August 2026?

The highest Elo 1 339 was achieved by GPT Image 2 from OpenAI. Second place is Reve 2.1 with 1 299 and third is MAI-Image-2.5 with 1 270. The score reflects the result of human preferences in pairwise comparisons, not technical quality or accuracy of prompt adherence.

What does the Elo difference between the GPT Image 2 model and the Recraft V4.1 Utility model mean?

The difference exceeds one hundred points, which in the Elo system indicates a strong preference of evaluators for GPT Image 2 outputs. Recraft V4.1 Utility placed last with 1 205. In practice this does not mean Recraft does not produce usable results, but that in direct comparison people more often choose the image from OpenAI.

Are the Nano Banana 2 Lite and GPT Image 1.5 models equal in the leaderboard?

Both models achieved an identical Elo of 1 263 and share fourth and fifth place. Nano Banana 2 Lite comes from Google, GPT Image 1.5 from OpenAI. The identical score indicates comparable likability in comparative tests, but does not imply similarity in style, speed, or prompt adherence.

DATA SOURCES AND METHOD

Všechna čísla v tomto srovnání pocházejí z uvedených zdrojů k datu 4 August 2026. Grafy generuje institut CIAD přímo ze zdrojových dat; textová analýza čísla nikdy nedopočítává ani neodhaduje.