Skip to content

Add ImageBench V1 to Section 6 (Benchmark vs. eval / integrity) - #71

Open
dh7 wants to merge 1 commit into
benchflow-ai:mainfrom
dh7:add-imagebench
Open

Add ImageBench V1 to Section 6 (Benchmark vs. eval / integrity)#71
dh7 wants to merge 1 commit into
benchflow-ai:mainfrom
dh7:add-imagebench

Conversation

@dh7

@dh7 dh7 commented Aug 1, 2026

Copy link
Copy Markdown

Adds one entry to Section 6 (Benchmark vs. eval / integrity).

Why it clears the bar:

  • Show your work: 40+ current text-to-image models scored on 192 prompts across 6 categories, with VLM judges + Estimated Preference Score, refreshed as new models ship. Every generated image is published rather than sampled, so failures are inspectable.
  • The angle for this section: "no-hidden-outputs" as a concrete integrity property for generative-image evals — the flip side of the cherry-picking that pervades T2I marketing. Complements the paper-heavy entries with a live worked example.
  • Open methodology + judge prompts: https://imagebench.ai/methodology-v1.

Happy to move it, reword the one-line note, or drop the 🆕 tag if you prefer.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant