神经网络在处理数字符号任务时,会自发形成类似符号的数值变量。
Emergent Symbol-like Number Variables in Artificial Neural Networks
- 通过神经子空间视角,将网络活动解释为简化符号算法。
- 不同架构下对符号算法的对齐程度从高度匹配到完全失败不等。
- 递归模型能生成分级的类符号数值变量,适合可解释性研究者阅读。
本文通过多种因果与理论方法,分析神经网络在基于序列的数字符号任务中如何表示数值信息。使用自回归的GRU、LSTM和Transformer,在仅依赖任务结构隐含数值信息的任务上训练模型。研究表明,当从个体神经元转向神经子空间视角时,可通过简化符号算法(SAs)有效解释原始网络活动。利用分布式对齐搜索(DAS),发现对齐程度随网络架构、维度和任务设定而变化,可能高度匹配、近似或完全失败。为此扩展了DAS框架,引入更灵活的线性对齐函数(LAFs),对比原有正交对齐函数(OAFs)。具体案例分析验证了对神经子空间的因果干预对可解释性的价值,并证明递归模型可在其活动中发展出分级的类符号数值变量。此外,浅层Transformer采用非马尔可夫解法——依赖非累积隐藏状态——且在注意力层不足时必须如此。
原文摘要 · Abstract (English)
What types of numeric representations emerge in neural systems, and what would a satisfying answer to this question look like? In this work, we interpret Neural Network (NN) solutions to sequence based number tasks using a variety of methods to understand how well we can interpret them through the lens of interpretable Symbolic Algorithms (SAs) -- precise programs describable by rules and typed, mutable variables. We use autoregressive GRUs, LSTMs, and Transformers trained on tasks where the correct tokens depend on numeric information only latent in the task structure. We show through multiple causal and theoretical methods that we can interpret raw NN activity through the lens of simplified SAs when we frame the activity in terms of neural subspaces rather than individual neurons. Using Distributed Alignment Search (DAS), we find that, depending on network architecture, dimensionality, and task specifications, alignments with SA's can be very high, or they can be only approximate, or fail altogether. We extend our analytic toolkit to address the failure cases by expanding the DAS framework to a broader class of alignment functions that more flexibly capture NN activity in terms of interpretable variables from SAs, and we provide theoretic and empirical explorations of Linear Alignment Functions (LAFs) in contrast to the preexisting Orthogonal Alignment Functions (OAFs). Through analyses of specific cases we confirm the usefulness of causal interventions on neural subspaces for NN interpretability, and we show that recurrent models can develop graded, symbol-like number variables in their neural activity. We further show that shallow Transformers learn very different solutions than recurrent networks, and we prove that such models must use anti-Markovian solutions -- solutions that do not rely on cumulative, Markovian hidden states -- in the absence of sufficient attention layers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。