arXiv:2603.20662cs.AIcs.CV2026-03被引 6

解析视觉语言模型中注意力头的空间推理功能,发现其稀疏且关键。

Attention in Space: Functional Roles of VLM Heads for Spatial Reasoning

  • 通过分解空间问题为分步子问题,识别不同注意力头的功能角色。
  • 空间相关注意力头数量少于其他认知功能,且激活后能提升推理准确率。
  • 适合关注多模态模型可解释性与空间推理优化的研究者。

尽管大型视觉语言模型(VLM)取得显著进展,空间推理仍是持续挑战。本文从机制可解释性角度,研究VLM中注意力头在空间推理中的功能角色。提出CogVSR数据集,将复杂空间推理问题分解为逐步子问题,模拟人类链式思维,每个子问题对应特定认知功能(如空间感知或关系推理)。基于此,构建探测框架以识别并表征特定功能的注意力头。跨多种VLM架构的分析显示,这些功能头普遍稀疏,数量和分布随功能而异。值得注意的是,空间专用头的数量少于其他认知功能头,凸显其稀缺性。提出方法激活潜在空间头,改善空间理解。干预实验进一步证明其关键作用:移除功能头导致性能下降,强调它们则提升准确率。本研究为理解VLM如何关注空间提供了新见解,并为增强多模态模型的复杂空间推理能力开辟路径。

原文摘要 · Abstract (English)

Despite remarkable advances in large Vision-Language Models (VLMs), spatial reasoning remains a persistent challenge. In this work, we investigate how attention heads within VLMs contribute to spatial reasoning by analyzing their functional roles through a mechanistic interpretability lens. We introduce CogVSR, a dataset that decomposes complex spatial reasoning questions into step-by-step subquestions designed to simulate human-like reasoning via a chain-of-thought paradigm, with each subquestion linked to specific cognitive functions such as spatial perception or relational reasoning. Building on CogVSR, we develop a probing framework to identify and characterize attention heads specialized for these functions. Our analysis across diverse VLM families reveals that these functional heads are universally sparse, vary in number and distribution across functions. Notably, spatially specialized heads are fewer than those for other cognitive functions, highlighting their scarcity. We propose methods to activate latent spatial heads, improving spatial understanding. Intervention experiments further demonstrate their critical role in spatial reasoning: removing functional heads leads to performance degradation, while emphasizing them enhances accuracy. This study provides new interpretability driven insights into how VLMs attend to space and paves the way for enhancing complex spatial reasoning in multimodal models.

空间推理注意力机制可解释性多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。