ImageBench

The only generative image benchmark that shows the images

36 models, 192 prompts, 6 categories — every output published. Judge with your own eyes which model is best for your use case, your budget, your quality bar.

The results at a glance

A preview of the top 10 across both benchmarks, plus how quality trades off against model size and price. Follow any link for the full, filterable results.

Benchmark V1

Ranked by Overall — capability (graded pass/fail across 192 prompts) blended with aesthetic preference.

See the full leaderboard
#
1openai/gpt-image-2
78.5
45.3s
2fal/google/nano-banana-2
73.0
28.1s
3fal/google/nano-banana-pro
66.2
23.4s
4fal/bytedance/seedream-v5-pro
65.6
147.4s
5local/boogu-image-turbo
61.9
9.9s
6bfl/flux-2-pro
60.1
11.8s
7bfl/flux-2-max
59.6
26.7s
8fal/bytedance/seedream-v4
59.4
14.1s
9local/qwen-image-2512-20b
58.6
80.2s
10local/flux-2-klein-9b
56.6
8.5s

Quality vs. size

Overall score against model size for the models we run locally — deeper green is the sweet spot.

smaller & higher quality203040506070800B4B8B12B16B20BModel size (billions of parameters)Overall scoreboogu-image-turbo — 10B, 61.9boogu-image-turboqwen-image-2512-20b — 20B, 58.6qwen-image-2512-20bflux-2-klein-9b — 9B, 56.6flux-2-klein-9bkrea-2-turbo — 12B, 56.5krea-2-turbokrea-2-turbo-no-filter — 12B, 53.8krea-2-turbo-no-filterflux-2-klein-4b — 4B, 51.5flux-2-klein-4bsefi-image-5b-base — 5B, 49.7sefi-image-5b-basez-image-turbo-6b — 6B, 49.2z-image-turbo-6bsefi-image-5b-rl — 5B, 47.7sefi-image-5b-rlz-image-6b — 6B, 47.4z-image-6bmage-flow-turbo — 4B, 47.3mage-flow-turbobonsai-image-ternary-4b — 4B, 46.5bonsai-image-ternary-4bprxpixel-t2i-7b — 7B, 46.3prxpixel-t2i-7bhidream-i1-full-17b — 17B, 42.9hidream-i1-full-17bkrea-2-raw — 12B, 42.1krea-2-rawsefi-image-5b-turbo — 5B, 41.9sefi-image-5b-turbonucleus-image-17b-a2b — 17B, 36.6nucleus-image-17b-a2bsefi-image-2b-turbo — 2B, 33.6sefi-image-2b-turbosana-1.5-1.6b — 1.6B, 33.0sana-1.5-1.6bSee it full-size

Quality vs. price

Overall score against API price per image for hosted models — deeper green is the sweet spot.

cheaper & higher quality20406080100$0.00$0.04$0.08$0.12$0.16API price per image (USD)Overall scoregoogle/nano-banana-2 — $0.080, 73.0google/nano-banana-2google/nano-banana-pro — $0.150, 66.2google/nano-banana-probytedance/seedream-v5-pro — $0.068, 65.6bytedance/seedream-v5-probfl/flux-2-pro — $0.030, 60.1bfl/flux-2-probfl/flux-2-max — $0.070, 59.6bfl/flux-2-maxbytedance/seedream-v4 — $0.030, 59.4bytedance/seedream-v4bfl/flux-2-klein-9b — $0.015, 55.3bfl/flux-2-klein-9bbytedance/seedream-v5-lite — $0.035, 52.7bytedance/seedream-v5-litebfl/flux-2-klein-4b — $0.014, 49.8bfl/flux-2-klein-4bideogram/v3 — $0.060, 48.5ideogram/v3bria/fast — $0.028, 47.9bria/fastkrea/v2-medium — $0.030, 46.9krea/v2-mediumkrea/v2-large — $0.060, 45.9krea/v2-largerecraft/v3 — $0.040, 45.5recraft/v3See it full-size

RealBench V1

Realism leaderboard — can a model's output pass as a real photograph? Scored by human votes.

See the full realism benchmark
#ModelRealism scoreRated real / votesImages
1fal/google/nano-banana-pro
53%
1,907 / 3,609141
2fal/bytedance/seedream-v4
47%
1,679 / 3,539139
3openai/gpt-image-2
46%
1,572 / 3,418139
4local/z-image-turbo-6b
46%
1,646 / 3,603140
5bfl/flux-2-max
41%
1,473 / 3,592141

Frequently asked questions