arXiv:2607.27069cs.CVcs.AI2026-07

提出视觉信用审计方法,检测多模态模型是否真依赖图像做空间推理。

Visual Credit Audit for Multimodal Spatial Reasoning

论文配图:Visual Credit Audit for Multimodal Spatial Reasoning
图 1 · 摘自论文原文
  • 通过对比图文、无图等控制条件,评估图像对答案的实际贡献
  • 发现12.73%-26.25%正确答案未获图像支持,存在虚假自信
  • 适合关注模型可解释性与评测可信度的研究者

封闭式空间推理基准可能在图像支持很弱时仍奖励正确答案。视觉信用审计(VCA)分离两个评估目标:一是图像是否比纯文本或空白条件为模型决策提供额外支持;二是模型是否响应特定关系的视觉证据。第一项审计无需训练或标签,也无需答案翻转。引入标签后得到依赖信用正确性(D-CC),在正确样本上等于同控组金标准正向增益;预测对齐则扩展至错误情形。在四个开放多模态大模型和两个空间基准上,12.73%-26.25%的决策虽正确却未获得图像信用。相同分割图像置换使D-CC下降21.25-47.80点,所有配对95%置信区间均高于零。固定像素关系对比及3×3证据源因子实验表明,零控制无法识别关系响应。在受控的正确但未获信用一致决策中,关系反转响应率达81.57%-100.00%,而总答案变更率仅为32.11%。对108个几何兼容编辑的独立审计结果提供了自然图像对应性的有界检验。因此,VCA将基准成功分解为正确性、额外图像支持与关系一致性响应三部分。

原文摘要 · Abstract (English)

Closed yes/no spatial benchmarks can reward a correct answer even when the image adds little support beyond no-image contexts. Under a fixed forced-choice interface, Visual Credit Audit (VCA) separates two estimands: whether the benchmark image gives the model's declared decision more support than text-only and blank controls, and whether the model responds to relation-specific visual evidence. The first audit is training- and label-free and does not require an answer flip. Applying labels yields dependence-credited correctness (D-CC); on correct items, it equals same-control gold-aligned positive gain, while prediction alignment extends the audit to errors. Across four open MLLMs and two spatial benchmarks, 12.73-26.25% of decisions are correct yet uncredited. Matched same-split image permutation reduces D-CC by 21.25-47.80 points, with every paired 95% interval above zero. Fixed-pixel relation contrasts and a 3x3 evidence-source factorial show why null controls cannot identify relation response. Among controlled correct-but-uncredited agreement decisions, response to relation reversal spans 81.57-100.00%, while 32.11% pooled change answer. Independently audited outcomes on 108 geometry-compatible edits provide a bounded natural-image correspondence check. VCA thereby decomposes benchmark success into correctness, additional image support, and relation-consistent response.

多模态空间推理模型审计可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。