揭秘视觉语言模型如何定位物体,发现关键计算路径。
Mechanisms of Object Localization in Vision-Language Models

- 通过消融与注意力分析,发现定位依赖物体对齐标记的容器化机制。
- 仅少数注意力头起关键作用,且在不同模型中分布层位不同。
- 结果适用于改进模型设计,尤其关注精准定位任务的研究者。
视觉-语言模型(VLMs)在关联视觉与文本信息方面表现优异,但在基础分类和定位任务上仍存在困难。尽管分类机制研究较多,定位过程仍不清晰。本文研究了两种代表性模型:LLaVA-1.5 和 InternVL-3.5,采用一系列机制可解释性工具,包括标记消融、注意力击穿和因果中介分析。结果显示,定位由一种‘容器化’机制驱动:物体对齐标记定义了物体的空间范围,而边界内标记的语义排列对预测框影响甚微。只有极少数注意力头对分类和定位具有因果影响,集中在 LLaVA 的早期-中期层,InternVL 则在中期-晚期层。两项任务共享部分早期处理,但最终依赖于大致独立的专用注意力头。本研究首次提供了 VLM 中定位的层级与头级解析,揭示了狭窄的计算路径,可为未来模型设计和对齐目标提供指导。
原文摘要 · Abstract (English)
Visually-grounded language models (VLMs) are highly effective in linking visual and textual information, yet they often struggle with basic classification and localization tasks. While classification mechanisms have been studied more extensively, the processes that support object localization remain poorly understood. In this work, we investigate two representative families, LLaVA-1.5 and InternVL-3.5, using a suite of mechanistic interpretability tools, including token ablations, attention knockout, and causal mediation analysis. We find that localization is driven by a containerization mechanism in which object-aligned tokens define the spatial extent of the object, while the semantic arrangement of tokens within those boundaries is largely irrelevant to the predicted box. Only a very small set of attention heads mediates the causal effect for both classification and localization, concentrating in early-mid layers for LLaVA and mid-late layers for InternVL. The two tasks share some early processing but ultimately depend on largely distinct specialized heads. Overall, we provide the first layer- and head-level account of localization in VLMs, revealing narrow computational pathways that can guide future model design and grounding objectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。