视觉语言模型常高估单模态信息的充分性,低估缺失模态恢复的影响。
I Don't Miss You, but I Do: Self-Explanation Faithfulness of Modality Missingness in Vision-Language Models

- 通过干预实验检验模型对模态缺失的自解释能力,对比预测与实际行为差异。
- 模型预测仅8.8%改变率,实际执行却达72.1%,62/64情况严重低估变化幅度。
- 自解释普遍不实,尤其在互补与同构数据中高估单模态作用,适合可信度研究者参考。
视觉语言模型在部分输入模态缺失的场景中应用日益广泛,但我们对其能否真实解释缺失信息如何影响自身预测知之甚少。本文提出一种干预评估协议:模型需声明各模态单独支持程度、恢复缺失模态是否改变判断、现有证据是否足够;随后执行相应模态干预,比较其预测与实际行为。我们在两个模型家族的八款开源VLM上,针对四类任务(互补与同构文本-图像设置、多视角驾驶场景)进行评估。结果发现,模型系统性高估可用模态证据的充分性。任务级预测变化率中位数仅为8.8%,而实际执行变化率高达72.1%,且62/64种情形中均存在严重低估。不足声明极少,但一旦出现则高度准确:恢复模态后答案改变的中位率达78%-100%。回溯性自解释也呈现相同趋势:在互补数据中高估单模态充分性,在同构数据中高估单一表征作用。结果表明,当前VLM系统性误判其预测对可用与缺失模态证据的依赖关系,凸显以可执行干预作为自解释行为基准的重要性。
原文摘要 · Abstract (English)
Vision-language models are increasingly used in settings where some input modalities may be unavailable, yet we know little about whether they can faithfully explain how such missing information affects their own predictions. We introduce an interventional protocol for evaluating self-explanations of modality dynamics: models state what each modality alone would support, whether restoring a missing modality would change their answer, and whether the available evidence is sufficient; we then execute the corresponding modality intervention and compare these claims with the model's realized behavior. We evaluate eight open-weight VLMs from two model families across four tasks spanning complementary and isomorphic text-image settings and a multi-view driving setting. We find a systematic tendency to overstate the sufficiency of available modality evidence. Models substantially underestimate the effect of restoring missing modalities: task-level median predicted change rates are at most 8.8%, while the corresponding executed change rates reach 72.1%, with underprediction in 62 of 64 model-task-condition settings. Insufficiency claims are rare, but precise when produced: restoring the modality changes the answer in a median of 78-100% of flagged cases. Retrospective self-explanations show the same tendency: on complementary data, models over-credit single-modality sufficiency; on isomorphic data, they over-credit single representation sufficiency relative to their executed behavior. Together, these results show that VLMs systematically mischaracterize how their predictions depend on available and missing modality evidence, motivating executable interventions as a behavioral ground truth for evaluating multimodal self-explanations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。