发现神经网络解模加法的通用抽象算法
Uncovering a Universal Abstract Algorithm for Modular Addition in Neural Networks
- 用近似中国剩余定理统一解释多种网络的模加法行为
- 深度网络只需O(log n)特征即可泛化,实验验证成立
- 适合研究模型可解释性与通用算法的学者
我们提出一个可检验的普遍性假设:看似不同的神经网络在简单模加法任务中表现出的解法,实则由同一抽象算法统一驱动。以往研究将神经元层级表征差异视为不同算法证据,但我们通过跨神经元、神经元簇和全网的多层分析,证明多层感知机与变压器均普遍实现一种称为近似中国剩余定理的抽象算法。关键在于引入近似陪集,并证明神经元仅在这些陪集上激活。此外,该理论适用于深度神经网络(DNNs),预测具有可训练嵌入或超过一层隐藏层的DNNs仅需O(log n)特征即可实现通用学习,实验结果予以验证。本工作首次提供基于理论的多层网络模加法求解解释,推动可泛化可解释性,并为群乘法等任务提出可检验的普遍性假说。
原文摘要 · Abstract (English)
We propose a testable universality hypothesis, asserting that seemingly disparate neural network solutions observed in the simple task of modular addition are unified under a common abstract algorithm. While prior work interpreted variations in neuron-level representations as evidence for distinct algorithms, we demonstrate - through multi-level analyses spanning neurons, neuron clusters, and entire networks - that multilayer perceptrons and transformers universally implement the abstract algorithm we call the approximate Chinese Remainder Theorem. Crucially, we introduce approximate cosets and show that neurons activate exclusively on them. Furthermore, our theory works for deep neural networks (DNNs). It predicts that universally learned solutions in DNNs with trainable embeddings or more than one hidden layer require only O(log n) features, a result we empirically confirm. This work thus provides the first theory-backed interpretation of multilayer networks solving modular addition. It advances generalizable interpretability and opens a testable universality hypothesis for group multiplication beyond modular addition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。