视觉语言模型评判时常忽略图像,偏好信息量大的答案。
When Vision-Language Models Judge Without Seeing: Exposing Informativeness Bias

- 先修正答案与图像的不一致,再对比判断
- 减少17%的信息量偏差,提升9.8%性能
- 适合评估视觉语言模型可靠性的研究者
VLM-as-a-Judge在自动评估视觉语言模型(VLMs)时的可靠性至关重要。尽管近期取得进展,我们的分析发现,VLM-as-a-Judge在决策时常忽略图像内容,反而盲目青睐信息量更大的答案,即使该答案与图像内容矛盾。我们称此现象为信息量偏差(informativeness bias),严重削弱了评判可靠性。为此,我们提出BIRCH(Balanced Informativeness and CoRrectness with a Truthful AnCHor)评判范式:首先修正候选答案与图像内容的不一致性,再基于修正后的版本进行对比。该方法将评判焦点从信息量转向图像对齐的正确性。在多个模型和基准上的实验表明,BIRCH将信息量偏差降低最多17%,性能提升达9.8%。本工作揭示了当前VLM-as-a-Judge系统中一个被忽视但根本性的缺陷,强调需采用更严谨的设计原则。
原文摘要 · Abstract (English)
The reliability of VLM-as-a-Judge is critical for the automatic evaluation of vision-language models (VLMs). Despite recent progress, our analysis reveals that VLM-as-a-Judge often pays limited attention to the image when making decisions. Instead, they often blindly favor the more informative answer, even when they can recognize it conflicts with the image content. We call this problem informativeness bias, which significantly undermines judge reliability. To address it, we propose BIRCH (Balanced Informativeness and CoRrectness with a Truthful AnCHor), a judging paradigm that first corrects inconsistencies with the image content in candidate answers, and then compares the answers against this corrected version. This shifts the judge's focus from informativeness to image-grounded correctness. Experiments on multiple models and benchmarks show that BIRCH reduces informativeness bias by up to 17%, resulting in performance gains of up to 9.8%. Our work reveals an overlooked but fundamental flaw in current VLM-as-a-Judge systems and highlights the need for more principled designs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。