神经网络的向量表示其实暗含符号结构,让黑箱模型更可解释。
The Emergent Symbolic Structure of Artificial Neural Networks
- 用符号结构替代神经网络的向量表示,行为几乎不变。
- 小网络和大语言模型在算术、逻辑等任务中均能被符号近似。
- 通过干预符号结构可精准调控大模型行为,适合可解释性研究者。
现代人工智能系统在看似不适宜的领域表现优异。传统认知科学认为智能基于符号组合(如逻辑公式),但当前最强的AI系统依赖神经网络,以连续向量表示信息。尽管向量看似无法捕捉语言、逻辑等结构化内容,神经网络却在这些领域表现出色。本文提出:神经网络的内部表征可能隐含符号结构。实验证明,多种神经网络的向量表示可被符号结构精确近似——用闭式方程替代整个表征生成过程后,模型行为基本不变。该现象适用于小规模列表操作网络及大语言模型(LLMs)在算术、逻辑、代码与语言四大核心符号任务中的表现。进一步地,对符号结构的精确干预可实现对LLM行为的定向修改,证明其行为依赖于所识别的符号结构。这项工作为调和符号主义与向量主义的长期分歧提供了新路径。
原文摘要 · Abstract (English)
Modern systems in artificial intelligence (AI) somehow excel in domains for which they seem poorly suited. Intelligence has traditionally been modeled as operating over structured combinations of symbols, such as logical formulas. However, the strongest modern AI systems are based on neural networks, which instead represent information in continuous vectors. Vectors seem inadequate for capturing the structure of language, logic, and other cognitive domains, yet neural networks achieve impressive performance in these areas. How do they do it? In this work, we propose a potential answer: Despite appearances, perhaps the internal representations of neural networks implicitly realize symbolic structure. In support of this hypothesis, we show that the vector representations of a variety of neural networks can be closely approximated with symbolic structures: we can replace the network's entire representation-generating process with a closed-form equation instantiating a symbolic structure, and the network's behavior remains largely unchanged. This finding holds for both small-scale neural networks trained to manipulate lists as well as large language models (LLMs) operating in four domains that are central in symbolic traditions: arithmetic, logic, computer code, and language. Further, our symbolic approximation allows us to modify an LLM's behavior in targeted ways via precise interventions on its internal representations, showing that the LLM's behavior is reliant on the symbolic structures we have identified. This work provides a potential way to reconcile longstanding symbolic conceptions of intelligence with the vector-based nature of modern AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。