arXiv:2510.13796cs.CLcs.CV2025-10被引 5

发现语言模型中符号意义通过注意力机制自发形成

The Mechanistic Emergence of Symbol Grounding in Language Models

  • 通过因果分析追踪中间层注意力头聚合环境信息
  • 符号接地集中在中间层,由注意力头整合环境信号
  • 适用于多模态对话与多种架构,不适用于单向LSTM

符号接地(Harnad, 1990)指词等符号如何通过与真实世界感官运动经验的关联获得意义。近期研究显示,在无显式接地目标的情况下,大规模(视觉-)语言模型中可能已出现接地现象。然而,其具体发生位置及驱动机制仍不清楚。为此,我们提出一个受控评估框架,系统追踪内部计算中接地现象的生成路径,采用机制与因果分析方法。结果表明,接地集中于中间层计算,通过注意力头聚合环境信号以支持语言形式预测。该现象在多模态对话和不同架构(Transformer与状态空间模型)中均可复现,但在单向LSTM中未出现。研究提供了行为与机制双重证据,证明符号接地可在语言模型中自发产生,对预测并可能控制生成可靠性具有实际意义。

原文摘要 · Abstract (English)

Symbol grounding (Harnad, 1990) describes how symbols such as words acquire their meanings by connecting to real-world sensorimotor experiences. Recent work has shown preliminary evidence that grounding may emerge in (vision-)language models trained at scale without using explicit grounding objectives. Yet, the specific loci of this emergence and the mechanisms that drive it remain largely unexplored. To address this problem, we introduce a controlled evaluation framework that systematically traces how symbol grounding arises within the internal computations through mechanistic and causal analysis. Our findings show that grounding concentrates in middle-layer computations and is implemented through the aggregate mechanism, where attention heads aggregate the environmental ground to support the prediction of linguistic forms. This phenomenon replicates in multimodal dialogue and across architectures (Transformers and state-space models), but not in unidirectional LSTMs. Our results provide behavioral and mechanistic evidence that symbol grounding can emerge in language models, with practical implications for predicting and potentially controlling the reliability of generation.

符号接地注意力机制语言模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。