arXiv:2604.21523cs.CVcs.CL2026-04

发现评估模型对生成缺陷视而不见,可能误导AI质量判断。

Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Models

论文配图:Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Models
图 1 · 摘自论文原文
  • 用针对性干扰测试评估模型对错误的识别能力
  • 超50%的错误未被发现,尤其难识别空间与组合错误
  • 适合关注AI评估可靠性的研究者与开发者参考

大型视觉语言模型(VLM)正被广泛用于评估图像到文本(I2T)和文本到图像(T2I)生成任务的输出质量。然而其评估可靠性尚未充分研究。本文系统评估了4种主流VLM在两类任务中的表现,引入针对物体幻觉、空间推理、事实一致性与视觉保真度等关键错误维度的定向扰动。基于超过4000个扰动实例、覆盖40种扰动类型的全面基准,采用单答案评分、成对比较与参考引导三种范式进行测试。结果表明:当前评估型VLM存在显著盲区——在某些情况下超过50%的劣化输出未被识别,尤其难以察觉细粒度的组合与空间错误,且对违背输入图像内容的幻觉信息反应迟钝。成对比较相对更可靠,但失败率依然存在。这些发现揭示了当前评估VLM的不可靠性,警示其在基准测试与开发决策中的使用需谨慎。代码与数据已公开。

原文摘要 · Abstract (English)

Large Vision-Language Models (VLMs) are increasingly used to evaluate outputs of other models, for image-to-text (I2T) tasks such as visual question answering, and text-to-image (T2I) generation tasks. Despite this growing reliance, the reliability of these Evaluator VLMs remains under explored. In this work, we systematically evaluate the reliability of Evaluator VLMs across both I2T and T2I tasks. We introduce targeted perturbations that degrade output quality along key error dimensions, including object hallucinations, spatial reasoning, factual grounding, and visual fidelity. These perturbations test whether Evaluator VLMs can reliably account for these quality degrading errors in their evaluations. Using a comprehensive benchmark of over 4000 perturbed instances spanning 40 perturbation dimensions, we evaluate 4 prominent VLMs using single-answer scoring, pairwise comparison, and reference-guided paradigms. Our findings reveal that current VLM evaluators exhibit substantial blind spots: they often fail to detect perturbed outputs - in some cases exceeding 50%, struggle particularly with fine-grained compositional and spatial errors, and are often insensitive to hallucinated content that contradicts the input image. Pairwise comparison proves more reliable, though failure rates persist. These results highlight the unreliable nature of current Evaluator VLMs and urge caution in their deployment for benchmarking and development decisions. Code and data have been made publicly available.

视觉语言模型评估可靠性幻觉检测AI评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。