arXiv:2602.09411cs.CV2026-02被引 2

用修正版视觉语言模型高效评估图像生成质量,比人工投票快且准。

K-Sort Eval: Efficient Preference Evaluation for Visual Generation via Corrected VLM-as-a-Judge

  • 基于人类投票数据训练修正模型,提升VLM评分与真人偏好一致性。
  • 动态匹配策略让每次评估只用不到90次模型运行,效率大幅提升。
  • 适合需要快速、低成本对比多个图像生成模型的研究者使用。

视觉生成模型的快速发展催生了对更可扩展、符合人类偏好的评估方法的需求。尽管众包平台通过收集人类投票提供偏好评估,但成本高、耗时长,难以规模化。利用视觉语言模型(VLM)替代人工判断是一种有前景的方案,但其固有的幻觉和偏差会削弱与人类偏好的一致性,影响评估可靠性。此外,静态评估方式效率低下。本文提出K-Sort Eval,一种结合后验修正与动态匹配的可靠高效VLM评估框架。我们从K-Sort Arena中收集数千条人类投票数据,构建高质量数据集,每条包含K个模型的输出及其排名。新模型评估时,与已有模型进行(K+1)路自由对抗比较,VLM给出排名。为提升对齐度与可靠性,提出后验修正方法,基于VLM预测与人类标注的一致性,自适应调整贝叶斯更新中的后验概率。同时提出动态匹配策略,平衡不确定性与多样性,最大化每次比较的预期收益,从而实现更高效率。大量实验表明,K-Sort Eval的评估结果与K-Sort Arena高度一致,通常仅需少于90次模型运行,验证了其高效性与可靠性。

原文摘要 · Abstract (English)

The rapid development of visual generative models raises the need for more scalable and human-aligned evaluation methods. While the crowdsourced Arena platforms offer human preference assessments by collecting human votes, they are costly and time-consuming, inherently limiting their scalability. Leveraging vision-language model (VLMs) as substitutes for manual judgments presents a promising solution. However, the inherent hallucinations and biases of VLMs hinder alignment with human preferences, thus compromising evaluation reliability. Additionally, the static evaluation approach lead to low efficiency. In this paper, we propose K-Sort Eval, a reliable and efficient VLM-based evaluation framework that integrates posterior correction and dynamic matching. Specifically, we curate a high-quality dataset from thousands of human votes in K-Sort Arena, with each instance containing the outputs and rankings of K models. When evaluating a new model, it undergoes (K+1)-wise free-for-all comparisons with existing models, and the VLM provide the rankings. To enhance alignment and reliability, we propose a posterior correction method, which adaptively corrects the posterior probability in Bayesian updating based on the consistency between the VLM prediction and human supervision. Moreover, we propose a dynamic matching strategy, which balances uncertainty and diversity to maximize the expected benefit of each comparison, thus ensuring more efficient evaluation. Extensive experiments show that K-Sort Eval delivers evaluation results consistent with K-Sort Arena, typically requiring fewer than 90 model runs, demonstrating both its efficiency and reliability.

图像生成评估方法VLM高效评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。