研究发现视觉干扰能让LVLM误判图文匹配度,导致评分虚高。
Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation
- 设计多种视觉干扰,测试LVLM对图文一致性的判断
- 所有测试的LVLM在干扰下评分均被系统性抬高
- 即使用提示词缓解,偏差仍存在,适合评估者警惕
近期,大视觉语言模型(LVLM)已成为评判图文对齐性的首选工具,但其在视觉模态上的鲁棒性尚未得到充分探索。本文首次探讨关键问题:对抗性视觉扰动能否系统性地误导LVLM裁判,使其给出不公正的高分?我们定义了图像引发的潜在偏差,并分析其对LVLM裁判评价的影响。进一步提出一个细粒度、多领域元评估基准FRAME,其刻意构建了多样化的评分分布。将所定义的偏差引入该基准后,发现所有测试的LVLM裁判在各领域均表现出脆弱性,对干扰图像的评分持续偏高。深入分析表明,多种偏差叠加会增强其效应,成对评估也同样易受攻击。此外,基于提示的缓解策略无法消除视觉偏差,凸显当前LVLM评估体系的脆弱性,亟需更鲁棒的裁判机制。
原文摘要 · Abstract (English)
Recently, large vision-language models (LVLMs) have emerged as the preferred tools for judging text-image alignment, yet their robustness along the visual modality remains underexplored. This work is the first study to address a key research question: Can adversarial visual manipulations systematically fool LVLM judges into assigning unfairly inflated scores? We define potential image induced biases within the context of T2I evaluation and examine how these biases affect the evaluations of LVLM judges. Moreover, we introduce a novel, fine-grained, multi-domain meta-evaluation benchmark named FRAME, which is deliberately constructed to exhibit diverse score distributions. By introducing the defined biases into the benchmark, we reveal that all tested LVLM judges exhibit vulnerability across all domains, consistently inflating scores for manipulated images. Further analysis reveals that combining multiple biases amplifies their effects, and pairwise evaluations are similarly susceptible. Moreover, we observe that visual biases persist under prompt-based mitigation strategies, highlighting the vulnerability of current LVLM evaluation systems and underscoring the urgent need for more robust LVLM judges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。