发现视觉语言模型中专注推理的注意力头,揭示其稀疏分布与关键作用。
Investigating The Functional Roles of Attention Heads in Vision Language Models: Evidence for Reasoning Modules
- 构建认知分解数据集CogVision,模拟人类思维链分析注意力机制
- 识别出少数专用注意力头,其移除导致模型性能显著下降
- 为设计更类人感知与推理能力的模型提供新方向
尽管在多模态基准上表现优异,视觉语言模型(VLMs)仍被视为黑箱。本文提出一种新型可解释性框架,系统分析VLM内部机制,聚焦注意力头在多模态推理中的功能角色。为此,我们引入CogVision数据集,将复杂多模态问题分解为分步子问题,模拟人类思维链,每个子问题对应特定感知或认知功能,如高层视觉接收与推理。基于探针方法,我们识别出承担这些功能的注意力头,并将其定义为功能头。跨多种VLM家族的分析表明,这些功能头普遍稀疏,数量和分布随功能变化,并介导交互与层级组织。干预实验进一步证明其关键作用:移除功能头导致性能下降,强调它们则提升准确率。研究为理解VLM的认知结构提供了新视角,指明了构建更具人类对齐感知与推理能力模型的潜在路径。
原文摘要 · Abstract (English)
Despite excelling on multimodal benchmarks, vision-language models (VLMs) largely remain a black box. In this paper, we propose a novel interpretability framework to systematically analyze the internal mechanisms of VLMs, focusing on the functional roles of attention heads in multimodal reasoning. To this end, we introduce CogVision, a dataset that decomposes complex multimodal questions into step-by-step subquestions designed to simulate human reasoning through a chain-of-thought paradigm, with each subquestion associated with specific receptive or cognitive functions such as high-level visual reception and inference. Using a probing-based methodology, we identify attention heads that specialize in these functions and characterize them as functional heads. Our analysis across diverse VLM families reveals that these functional heads are universally sparse, vary in number and distribution across functions, and mediate interactions and hierarchical organization. Furthermore, intervention experiments demonstrate their critical role in multimodal reasoning: removing functional heads leads to performance degradation, while emphasizing them enhances accuracy. These findings provide new insights into the cognitive organization of VLMs and suggest promising directions for designing models with more human-aligned perceptual and reasoning abilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。