arXiv:2512.16853cs.CVcs.AI2025-12被引 32

旧评测基准过时,新基准GenEval 2更难且更贴近人类判断。

GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation

  • 用视觉基础概念和复合语义提升评测覆盖度
  • 旧基准误差达17.7%,新基准对当前模型更具挑战性
  • 引入软判别法提升与人类判断一致性,减少漂移

自动化文本到图像(T2I)模型评估困难,需依赖评分模型判断准确性,且测试提示必须对当前模型有挑战但对评分模型不能太难。我们指出,满足这些条件会导致基准漂移——静态基准无法跟上新模型能力。以最流行的T2I评测之一GenEval为例,其初始与人类判断高度一致,但随时间推移已严重偏离,当前模型误差高达17.7%。大规模人工研究证实该基准已饱和。为此,我们提出新基准GenEval 2,增强基础视觉概念覆盖与组合复杂度,对现有模型更具挑战。同时引入Soft-TIFA方法,通过整合视觉基元判断,相比整体评分器如VQAScore,更贴近人类判断且更不易漂移。尽管期望GenEval 2能长期有效,但避免漂移仍需持续审计与优化,凸显对T2I等自动评测基准的动态维护重要性。

原文摘要 · Abstract (English)

Automating Text-to-Image (T2I) model evaluation is challenging; a judge model must be used to score correctness, and test prompts must be selected to be challenging for current T2I models but not the judge. We argue that satisfying these constraints can lead to benchmark drift over time, where the static benchmark judges fail to keep up with newer model capabilities. We show that benchmark drift is a significant problem for GenEval, one of the most popular T2I benchmarks. Although GenEval was well-aligned with human judgment at the time of its release, it has drifted far from human judgment over time -- resulting in an absolute error of as much as 17.7% for current models. This level of drift strongly suggests that GenEval has been saturated for some time, as we verify via a large-scale human study. To help fill this benchmarking gap, we introduce a new benchmark, GenEval 2, with improved coverage of primitive visual concepts and higher degrees of compositionality, which we show is more challenging for current models. We also introduce Soft-TIFA, an evaluation method for GenEval 2 that combines judgments for visual primitives, which we show is more well-aligned with human judgment and argue is less likely to drift from human-alignment over time (as compared to more holistic judges such as VQAScore). Although we hope GenEval 2 will provide a strong benchmark for many years, avoiding benchmark drift is far from guaranteed and our work, more generally, highlights the importance of continual audits and improvement for T2I and related automated model evaluation benchmarks.

图像生成评测基准基准漂移人类对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。