arXiv:2603.27070cs.CV2026-03

通过神经元关联图分析视觉语言模型的结构,发现深层存在关键枢纽神经元。

Structural Graph Probing of Vision-Language Models

  • 用神经元共激活构建每层相关图,研究模型内部拓扑结构
  • 深层逐渐形成紧凑的跨模态枢纽神经元集群,干预可显著改变输出
  • 该方法介于局部归因与全电路恢复之间,适合解释多模态行为

视觉语言模型(VLMs)在多模态任务中表现优异,但其神经元群体间的计算组织方式仍不清晰。本文从神经拓扑角度出发,将每一层表示为由神经元-神经元共激活构建的层内相关图。这一视角使我们能够探究群体结构是否具有行为意义、如何随模态和深度变化,以及是否能识别出干预下的因果关键组件。结果表明,相关拓扑包含可恢复的行为信号;跨模态结构随深度逐步凝聚为一组紧凑的重复枢纽神经元,对其针对性扰动会显著改变模型输出。神经拓扑因此成为视觉语言模型可解释性的重要中间尺度:比局部归因更丰富,比完整电路恢复更可行,并且与多模态行为有实证关联。代码已公开于 https://github.com/he-h/vlm-graph-probing。

原文摘要 · Abstract (English)

Vision-language models (VLMs) achieve strong multimodal performance, yet how computation is organized across populations of neurons remains poorly understood. In this work, we study VLMs through the lens of neural topology, representing each layer as a within-layer correlation graph derived from neuron-neuron co-activations. This view allows us to ask whether population-level structure is behaviorally meaningful, how it changes across modalities and depth, and whether it identifies causally influential internal components under intervention. We show that correlation topology carries recoverable behavioral signal; moreover, cross-modal structure progressively consolidates with depth around a compact set of recurrent hub neurons, whose targeted perturbation substantially alters model output. Neural topology thus emerges as a meaningful intermediate scale for VLM interpretability: richer than local attribution, more tractable than full circuit recovery, and empirically tied to multimodal behavior. Code is publicly available at https://github.com/he-h/vlm-graph-probing.

视觉语言模型神经拓扑可解释性枢纽神经元

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。