用配对比较提升图像生成评估可靠性,开源模型表现超越大厂闭源模型
GenArena: How Can We Achieve Human-Aligned Evaluation for Visual Generation Tasks?
- 采用配对比较替代传统打分法,提升评估稳定性
- 评估准确率提升超20%,与权威榜单相关性达0.86
- 适合需要公平、自动化评估的视觉生成研究者
视觉生成模型的快速进步已超出传统评估方法的能力,促使采用视觉语言模型作为代理评判者。本文系统研究了广泛视觉生成任务中主流的绝对单点评分标准的可靠性。分析表明,该范式受限于随机不一致性及与人类感知的低对齐度。为此,我们提出GenArena统一评估框架,采用配对比较机制,实现稳定且与人类对齐的评估。关键发现:仅采用此配对协议,即可使现成开源模型表现超越顶级专有模型。实验显示,该方法使评估准确率提升超20%,与权威LMArena榜单的斯皮尔曼相关系数达0.86,远超点对点方法的0.36。基于GenArena,我们在多样任务上基准测试了先进视觉生成模型,为社区提供严谨、自动化的评估标准。
原文摘要 · Abstract (English)
The rapid advancement of visual generation models has outpaced traditional evaluation approaches, necessitating the adoption of Vision-Language Models as surrogate judges. In this work, we systematically investigate the reliability of the prevailing absolute pointwise scoring standard, across a wide spectrum of visual generation tasks. Our analysis reveals that this paradigm is limited due to stochastic inconsistency and poor alignment with human perception. To resolve these limitations, we introduce GenArena, a unified evaluation framework that leverages a pairwise comparison paradigm to ensure stable and human-aligned evaluation. Crucially, our experiments uncover a transformative finding that simply adopting this pairwise protocol enables off-the-shelf open-source models to outperform top-tier proprietary models. Notably, our method boosts evaluation accuracy by over 20% and achieves a Spearman correlation of 0.86 with the authoritative LMArena leaderboard, drastically surpassing the 0.36 correlation of pointwise methods. Based on GenArena, we benchmark state-of-the-art visual generation models across diverse tasks, providing the community with a rigorous and automated evaluation standard for visual generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。