测试视觉模型能否突破先验偏见,用工具增强的视觉证据是否有效。
Seeing What Is Actually There: PriVE-Bench and PriVE-Tools for Counterfactual Evaluation of Agentic Visual Evidence in VLMs

- 设计反事实图像对,区分模型是看图还是靠经验回答。
- 工具生成的局部视觉证据可提升部分模型的准确率。
- 即使有明确图像证据,仍有多数模型坚持语言先验,适合研究模型可靠性者阅读。
视觉-语言模型(VLMs)常依赖语言和类别先验回答视觉问题,而非基于图像内容。反事实图像提供自然诊断场景:当可见证据与常识冲突时,真正基于图像的模型应依据像素作答,而依赖先验的模型则输出符合惯例但视觉错误的答案。现有基准仅检验先验行为是否存在。本文进一步探究在工具增强与代理式视觉系统兴起背景下:额外的视觉证据视图能否帮助模型克服先验偏差?我们提出PriVE-Bench,一个通过原始与反事实图像对区分视觉接地回答与先验一致错误的基准。同时引入PriVE-Tools,一种受控的代理视觉扩展,评估边界框、裁剪、缩放面板和轮廓等工具生成的视觉证据是否改善反事实情境下的接地性。在开源与闭源VLMs上,对比原始输入、图像对输入及工具条件输入下的准确率、先验错误率与他回应率。结果表明,在模型能有效利用局部证据时,视觉工具有助于提升表现,但并非万能:多个模型即便面对明确视觉证据,仍持续遵循语言与类别先验。
原文摘要 · Abstract (English)
Vision-language models (VLMs) often answer visual questions using learned language and category priors rather than grounding their predictions in the image itself. Counterfactual images provide a natural diagnostic setting for this failure mode: when visible evidence contradicts what is usually true, a grounded model should answer from the pixels, while a prior-following model will produce a canonical but visually incorrect response. However, existing counterfactual benchmarks mainly ask whether such prior-following behavior exists. In this paper, we ask a further question motivated by the rise of tool-augmented and agentic vision systems: can additional visual evidence views help VLMs reason against their priors? We introduce PriVE-Bench, a Prior-vs-Visual Evidence Benchmark that uses paired original and counterfactual images to distinguish visually grounded answers from prior-consistent errors. We further introduce PriVE-Tools, a controlled agentic-vision-inspired extension that evaluates whether tool-derived visual evidence -- including bounding boxes, crops, zoom panels, and contours -- improves grounding under the same counterfactual conflicts. Across open- and closed-source VLMs, we compare raw, paired-image, and tool-conditioned inputs using accuracy, prior-following error rate, and other-response rate. Our results show that visual evidence tools can help in some settings, especially when models can use localized evidence effectively, but they are not a universal remedy: several models continue to follow language and category priors even when relevant visual evidence is explicitly provided.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。