拆解视觉语言模型的冲突检测与解决机制,提升模型可解释性。
Challenges in Understanding Modality Conflict in Vision-Language Models
- 通过线性探测和注意力模式分析,分离冲突检测与解决过程。
- 发现中间层存在可线性解码的冲突信号,且两类机制在不同阶段分化。
- 为提升复杂多模态场景下的模型鲁棒性提供可干预的解释路径。
本文指出,将冲突检测与冲突解决从视觉语言模型(VLM)中分离是当前的关键挑战,并提出潜在方法:利用线性探测的监督指标以及基于分组的注意力模式分析。我们对当前最先进的模型LLaVA-OV-7B进行了机制性研究,该模型在面对多模态冲突输入时表现出多样化的解决行为。结果表明,模型中间层存在可线性解码的冲突信号,且与冲突检测相关的注意力模式与解决模式在网络不同阶段出现分化。这些发现支持了检测与解决是功能上独立机制的假设。我们进一步讨论这种分解如何促进更可操作的可解释性分析,并实现针对模型鲁棒性的精准干预。
原文摘要 · Abstract (English)
This paper highlights the challenge of decomposing conflict detection from conflict resolution in Vision-Language Models (VLMs) and presents potential approaches, including using a supervised metric via linear probes and group-based attention pattern analysis. We conduct a mechanistic investigation of LLaVA-OV-7B, a state-of-the-art VLM that exhibits diverse resolution behaviors when faced with conflicting multimodal inputs. Our results show that a linearly decodable conflict signal emerges in the model's intermediate layers and that attention patterns associated with conflict detection and resolution diverge at different stages of the network. These findings support the hypothesis that detection and resolution are functionally distinct mechanisms. We discuss how such decomposition enables more actionable interpretability and targeted interventions for improving model robustness in challenging multimodal settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。