通过激活操控揭示多模态大模型的视觉表征机制。
Causal Probing for Internal Visual Representations in Multimodal Large Language Models

- 用激活引导方法主动探测内部视觉表征
- 实体概念局部化,抽象概念全局分布
- 适合研究模型内部机制与推理缺陷的学者
尽管多模态大语言模型(MLLMs)在多种任务中表现卓越,但其如何编码和锚定不同视觉概念的内部机制仍不明确。为此,我们提出一种基于激活引导的因果框架,主动探测并操纵内部视觉表征。在四个视觉概念类别上系统干预后,结果揭示概念编码存在分化:实体知识具有显著局部化特征,而抽象概念则全局分布于网络中。关键发现是,这种分化揭示了缩放定律的机制根源:增加模型深度对编码分布式复杂抽象概念至关重要,而实体始终高度局部化。此外,反向引导显示,阻断显式输出会引发潜在激活激增,暴露感知与生成间的补偿机制。最后,将分析扩展至视觉推理,发现感知与推理间存在脱节:虽能识别几何关系,但仅将其视为静态视觉特征,未能触发解决问题所需的程序性执行。
原文摘要 · Abstract (English)
Despite the remarkable success of Multimodal Large Language Models (MLLMs) across diverse tasks, the internal mechanisms governing how they encode and ground distinct visual concepts remain poorly understood. To unravel these mechanisms, we propose a causal framework based on activation steering to actively probe and manipulate internal visual representations. Through systematic intervention across four visual concept categories, our results reveal a divergence in concept encoding: entity knowledge is distinctively localized, whereas abstract concepts are globally distributed across the network. Critically, this divergence uncovers a mechanistic driver of scaling laws: increasing model depth is indispensable for encoding distributed and complex abstract concepts, whereas entities maintain a consistently high degree of localization. Furthermore, reverse steering uncovers that blocking explicit output triggers a surge in latent activations, exposing a compensatory mechanism between perception and generation. Finally, by extending our analysis to visual reasoning, we expose a disconnect between perception and reasoning: although MLLMs successfully recognize geometric relations, they treat them merely as static visual features, failing to trigger the procedural execution necessary for solving problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。