arXiv:2503.12992cs.AIcs.CL2025-03

发现语言模型神经元能基于激活水平识别语义片段,推动高层抽象形成。

Intra-neuronal attention within language models Relationships between activation and semantics

  • 通过激活区域定位神经元响应的语义片段,实现神经元内注意力机制。
  • 仅在高激活令牌上观察到激活与语义分段间的同态关系。
  • 该机制支持下层神经元进行语义重构,助力高层抽象生成。

本研究探讨了语言模型中感知机型神经元执行神经元内注意力的能力:即基于特定激活区域对它们高度响应的标记(tokens)所编码的合成思想类别中的同质语义段进行识别。研究目标是确定形式化神经元能否在激活划分与语义分段之间建立同态关系。结果表明,这种关系虽存在但较微弱,仅在高激活水平的标记层面成立。该神经元内注意力随后可促使下一层神经元进行语义重构,从而参与高层次语义抽象的逐步形成。

原文摘要 · Abstract (English)

This study investigates the ability of perceptron-type neurons in language models to perform intra-neuronal attention; that is, to identify different homogeneous categorical segments within the synthetic thought category they encode, based on a segmentation of specific activation zones for the tokens to which they are particularly responsive. The objective of this work is therefore to determine to what extent formal neurons can establish a homomorphic relationship between activation-based and categorical segmentations. The results suggest the existence of such a relationship, albeit tenuous, only at the level of tokens with very high activation levels. This intra-neuronal attention subsequently enables categorical restructuring processes at the level of neurons in the following layer, thereby contributing to the progressive formation of high-level categorical abstractions.

神经元机制语义抽象注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。