arXiv:2605.13047cs.CVcs.AI2026-05

用反事实方法揭示视觉大模型与人类在场景理解上的差距

Revealing the Gap in Human and VLM Scene Perception through Counterfactual Semantic Saliency

论文配图:Revealing the Gap in Human and VLM Scene Perception through Counterfactual Semantic Saliency
图 1 · 摘自论文原文
  • 通过因果消融测量物体对语义影响,实现黑盒模型的可解释性
  • 发现模型过度依赖大物体、中心物体和高显著性物体,忽略人物
  • 首次量化模型与人类在复杂场景理解中的系统性偏差差异

评估大型视觉语言模型(VLMs)在高层语义场景理解上是否与人类感知一致仍具挑战。传统白盒可解释性方法不适用于闭源架构,被动指标也无法分离因果特征。本文提出反事实语义显著性(CSS),一种黑盒、模型无关的框架,通过测量物体从场景中被因果消融后引发的语义变化,量化其重要性。为评估人工智能与人类的语义对齐程度,我们在包含307个复杂自然场景和1,306个高保真反事实变体的基准上,对比了主流VLMs与基于16,289次有效响应的人类心理物理学基线。分析显示存在普遍的场景理解差距:模型相对于人类更依赖大物体(尺寸偏差)、图像中心物体(中心偏差)和高显著性物体;而对图像中人物的依赖度低于人类参与者。模型的尺寸偏差是解释模型-人类语义差异的主要驱动因素。代码与数据将公开于https://github.com/starsky77/Counterfactual-Semantic-Saliency。

原文摘要 · Abstract (English)

Evaluating whether large vision-language models (VLMs) align with human perception for high-level semantic scene comprehension remains a challenge. Traditional white-box interpretability methods are inapplicable to closed-source architectures and passive metrics fail to isolate causal features. We introduce Counterfactual Semantic Saliency (CSS). This black-box, model-agnostic framework quantifies the importance of objects by measuring the semantic shift induced by their causal ablation from a scene. To evaluate AI-human semantic alignment, we tested prominent VLMs against a human psychophysics baseline comprising 16,289 valid responses across 307 complex natural scenes and 1,306 high-fidelity counterfactual variants. Our analysis reveals a pervasive scene comprehension gap: models exhibit an overreliance (relative to humans) on large objects (size bias), objects at the center of the image (center bias), and high saliency objects. In contrast, models rely less on people in the scenes than our human participants to describe the images. A model's size bias is a primary driver explaining variations in model-human semantic divergence. Code and data will be available at https://github.com/starsky77/Counterfactual-Semantic-Saliency.

视觉语言模型可解释性人类对齐反事实分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。