arXiv:2602.16849cs.LGmath.OC2026-02被引 4

解析神经网络如何学模加法,揭示频率、相位与彩票机制的作用。

On the Mechanism and Dynamics of Modular Addition: Fourier Features, Lottery Ticket, and Grokking

  • 通过相位对称与频率多样化,实现噪声抑制的集体决策。
  • 单个神经元学单一傅里叶频率,整体逼近模加逻辑的缺陷指示函数。
  • 解释了过参数化下特征涌现与模型泛化跃迁(grokking)的三阶段机制。

本文全面分析了两层神经网络学习模加任务的机制与训练动态。尽管已有研究发现单个神经元会学习特定频率的傅里叶特征并实现相位对齐,但未能解释这些特征如何组合成全局解。我们提出在过参数化训练中出现的分化条件,包含相位对称与频率多样化两个部分,证明其可使网络集体近似模加任务的错误指示函数。虽然单个神经元输出含噪信号,但相位对称支持多数投票机制,有效消除噪声,确保正确求和结果。此外,我们通过彩票机制解释随机初始化下特征的涌现:梯度流分析表明,每个神经元内频率竞争由初始谱幅值与相位对齐决定。技术上,我们严格刻画了逐层相位耦合动态,并利用微分包含比较引理形式化了竞争格局。最后,基于这些洞察,将grokking定性为三阶段过程:记忆阶段后接两次泛化阶段,由损失最小化与权重衰减的竞争驱动。

原文摘要 · Abstract (English)

We present a comprehensive analysis of how two-layer neural networks learn features to solve the modular addition task. Our work provides a full mechanistic interpretation of the learned model and a theoretical explanation of its training dynamics. While prior work has identified that individual neurons learn single-frequency Fourier features and phase alignment, it does not fully explain how these features combine into a global solution. We bridge this gap by formalizing a diversification condition that emerges during training when overparametrized, consisting of two parts: phase symmetry and frequency diversification. We prove that these properties allow the network to collectively approximate a flawed indicator function on the correct logic for the modular addition task. While individual neurons produce noisy signals, the phase symmetry enables a majority-voting scheme that cancels out noise, allowing the network to robustly identify the correct sum. Furthermore, we explain the emergence of these features under random initialization via a lottery ticket mechanism. Our gradient flow analysis proves that frequencies compete within each neuron, with the "winner" determined by its initial spectral magnitude and phase alignment. From a technical standpoint, we provide a rigorous characterization of the layer-wise phase coupling dynamics and formalize the competitive landscape using the ODE comparison lemma. Finally, we use these insights to demystify grokking, characterizing it as a three-stage process involving memorization followed by two generalization phases, driven by the competition between loss minimization and weight decay.

神经网络机制傅里叶特征Grokking模加法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。