arXiv:2602.08713cs.CVcs.LG2026-02

首次揭示语言模型如何通过微调获得空间感知能力

Towards Understanding Multimodal Fine-Tuning: Spatial Features

  • 用分阶段模型差分法追踪多模态微调中的表征变化
  • 发现少量注意力头负责激活空间关系编码特征
  • 适合关注多模态模型可解释性与机制理解的研究者

当前视觉-语言模型(VLMs)通过将视觉编码器与预训练语言模型结合,在多种任务上取得优异表现,但语言模型在多模态训练中如何调整表示、视觉能力何时出现仍不清晰。本文首次进行机制分析,采用分阶段模型差分法,隔离多模态微调引入的表征变化,揭示语言模型如何‘学会看见’。我们首先识别出在微调过程中涌现或重构的视觉偏好特征;接着发现其中一小部分特征能稳定编码空间关系,通过控制空间提示验证;最后追溯这些特征的因果激活源于少数注意力头。结果表明,分阶段模型差分法可定位空间化多模态特征出现的时间与位置,同时阐明视觉接地如何重塑原本仅文本的特征表示。该方法提升了多模态训练的可解释性,为理解并优化预训练语言模型获取视觉能力的机制奠定基础。

原文摘要 · Abstract (English)

Contemporary Vision-Language Models (VLMs) achieve strong performance on a wide range of tasks by pairing a vision encoder with a pre-trained language model, fine-tuned for visual-text inputs. Yet despite these gains, it remains unclear how language backbone representations adapt during multimodal training and when vision-specific capabilities emerge. In this work, we present the first mechanistic analysis of VLM adaptation. Using stage-wise model diffing, a technique that isolates representational changes introduced during multimodal fine-tuning, we reveal how a language model learns to "see". We first identify vision-preferring features that emerge or reorient during fine-tuning. We then show that a selective subset of these features reliably encodes spatial relations, revealed through controlled shifts to spatial prompts. Finally, we trace the causal activation of these features to a small group of attention heads. Our findings show that stage-wise model diffing reveals when and where spatially grounded multimodal features arise. It also provides a clearer view of modality fusion by showing how visual grounding reshapes features that were previously text-only. This methodology enhances the interpretability of multimodal training and provides a foundation for understanding and refining how pretrained language models acquire vision-grounded capabilities.

多模态可解释性语言模型空间感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。