arXiv:2605.17826cs.CVcs.AI2026-05

测试视觉语言模型在图像与常识冲突时能否忠实于视觉证据。

CounterCount: A Diagnostic Framework for Counting Bias in Vision Language Models

论文配图:CounterCount: A Diagnostic Framework for Counting Bias in Vision Language Models
图 1 · 摘自论文原文
  • 设计配对图像对比框架,检验模型对计数的依赖是图像还是先验。
  • 发现模型在反事实场景中准确率下降,仍依赖物体固有认知。
  • 提出注意力重加权策略,最高提升8%的计数准确性。

视觉语言模型(VLMs)在多模态推理中表现优异,但其回答是否基于视觉证据,还是受语言和世界先验驱动尚不明确。计数任务提供了精确的测试基准:当视觉证据与常规物体知识冲突时,模型应依赖图像而非典型数量。我们提出CounterCount,一个用于诊断VLM反事实计数的框架,包含成对的事实与反事实图像、编辑过的计数相关属性、验证答案及局部证据标注。评估近期VLMs发现,其在事实图像上表现良好,但在反事实属性变化下持续退化,表明即使存在矛盾的视觉证据,模型仍依赖物体级先验。通过局部标注分析,我们发现失败并非仅因视觉证据缺失或模糊,而是模型对计数相关视觉标记的关注不足。为此,我们引入统一的推理时注意力调制策略,重新加权选定视觉标记,使多个VLM的反事实计数准确率最高提升8%。总体而言,CounterCount揭示了由先验驱动的计数错误,并为未来VLM设计提供诊断洞察。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) excel at multimodal reasoning, yet it remains unclear whether their answers are grounded in visual evidence or driven by learned language and world priors. Counting provides a precise testbed: when visual evidence conflicts with canonical object knowledge, a model must rely on the image rather than a prototypical count. We introduce CounterCount, a diagnostic framework for counterfactual counting in VLMs, consisting of paired factual and counterfactual images with edited count-relevant attributes, verified answers, and localized evidence annotations. Evaluating recent VLMs, we find strong performance on factual images but consistent degradation under counterfactual attribute changes, indicating reliance on object-level priors even when contradictory visual evidence is present. Using localized annotations, we show that these failures are not solely due to missing or ambiguous visual evidence, but to models underweighting attention to count-relevant visual tokens. We introduce a unified inference-time attention modulation strategy that reweights selected visual tokens, improving counterfactual counting accuracy by up to 8% across multiple VLMs. Overall, CounterCount exposes prior-driven counting failures and provides diagnostic insights for designing future VLMs.

视觉语言模型计数偏差注意力机制诊断框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。