通过细粒度因果追踪,提升视觉语言模型的可解释性并减少幻觉。
Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic Interpretability and Hallucination Mitigation
- 提出跨模态因果追踪框架,分析视觉与文本令牌在各层组件中的作用。
- 发现中间层最后令牌的多头注意力对跨模态信息聚合至关重要。
- 引入推理时干预方法,无需训练即可有效抑制幻觉,保持速度与性能。
尽管大视觉语言模型(LVLMs)取得了显著进展,其机制可解释性仍研究不足。现有分析覆盖面有限,未能全面考察视觉与文本令牌、模型组件及全部层数。为此,我们提出细粒度跨模态因果追踪(FCCT)框架,系统量化对视觉对象感知的因果影响。该分析覆盖所有视觉与文本令牌、多头自注意力(MHSA)、前馈网络(FFN)和隐藏状态,贯穿解码器全部层级。首次揭示:中层最后一个令牌的MHSA在跨模态信息聚合中起关键作用;而FFN呈现三阶段分层进程,负责视觉对象表征的存储与传递。基于此,我们提出中间表示注入(IRI),一种无需训练的推理时干预技术,通过精准干预特定组件与层级的跨模态表示,增强感知能力并缓解幻觉。在五个主流基准与多个LVLM上一致取得最优表现,同时保持推理速度与基础性能。
原文摘要 · Abstract (English)
Despite the remarkable advancements of Large Vision-Language Models (LVLMs), the mechanistic interpretability remains underexplored. Existing analyses are insufficiently comprehensive and lack examination covering visual and textual tokens, model components, and the full range of layers. This limitation restricts actionable insights to improve the faithfulness of model output and the development of downstream tasks, such as hallucination mitigation. To address this limitation, we introduce Fine-grained Cross-modal Causal Tracing (FCCT) framework, which systematically quantifies the causal effects on visual object perception. FCCT conducts fine-grained analysis covering the full range of visual and textual tokens, three core model components including multi-head self-attention (MHSA), feed-forward networks (FFNs), and hidden states, across all decoder layers. Our analysis is the first to demonstrate that MHSAs of the last token in middle layers play a critical role in aggregating cross-modal information, while FFNs exhibit a three-stage hierarchical progression for the storage and transfer of visual object representations. Building on these insights, we propose Intermediate Representation Injection (IRI), a training-free inference-time technique that reinforces visual object information flow by precisely intervening on cross-modal representations at specific components and layers, thereby enhancing perception and mitigating hallucination. Consistent improvements across five widely used benchmarks and LVLMs demonstrate IRI achieves state-of-the-art performance, while preserving inference speed and other foundational performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。