模块化加法任务中神经表示受任务内在几何约束,形成二维循环结构。
Beyond Neural Collapse: Task-Intrinsic Geometry Governs Neural Representations in Modular Arithmetic

- 分类器权重先形成二维等角配置,嵌入向量被约束在相同平面
- 嵌入生成等价于圆上相位对齐,解为模素数的单频特征
- 相比传统神经坍缩,该方案在复杂度与对称性间取得更优平衡
尽管神经坍缩(NC)预测 $K$ 类均衡分类器应将终端表示组织为 $(K-1)$ 维等角紧框架(ETF),但模块化加法始终进入不同范式:网络压缩至二维循环几何,分类器权重与标记嵌入均位于圆周上。我们从三方面深化解释:首先,提出分层非均匀训练机制——下游分类器权重因密集交叉熵梯度驱动,在上游嵌入完全重组前即形成秩-2等角配置;一旦此分类平面确立,反向传播特征梯度将嵌入运动限制于同一平面,而权重衰减抑制正交分量。其次,该子空间锁定后,诱导的平面动力学可解释为 $S^1$ 上带熵正则化的输运过程;结合模加法标签,嵌入形成等价于相位对齐,其极小值为 $\/mathbb{Z}/P\/mathbb{Z}$ 的单频字符,即圆上等角点。第三,量化表明该解优于 NC:ETF 仅在交叉熵上获 $O(1)$ 优势,而循环秩-2 解在施坦纳或权重衰减代理下获得 $Θ(K)$ 优势,产生临界阈值 $λ_{\mathrm{crit}} = Θ(1/K)$。结果揭示分类器为何先行移动,嵌入如何随后对齐,说明模运算中的 '领悟' 不仅由最大分离决定,而是分离、对称与复杂度之间任务结构化的权衡。
原文摘要 · Abstract (English)
While neural collapse (NC) predicts that a $K$-class-balanced classifier should organize terminal representations as a $(K-1)$-dimensional simplex equiangular tight frame (ETF), modular addition consistently enters a different regime: networks compress to a two-dimensional cyclic geometry in which both classifier weights and token embeddings lie on circles. We refine the explanation of this phenomenon in three directions. First, we formalize a layerwise non-uniform training mechanism: downstream classifier weights are driven by dense cross-entropy gradients into a rank-2 equiangular configuration before upstream embeddings fully reorganize, and once this classifier plane forms, backpropagated feature gradients constrain embedding motion to the same plane while weight decay suppresses orthogonal components. Second, after this subspace locking, the induced in-plane dynamics admit an entropy-regularized transport interpretation on $S^1$; combined with modular-addition labels, this reduces embedding formation to phase alignment, whose minimizers are single-frequency characters of $\mathbb{Z}/P\mathbb{Z}$ and hence equal-angle points on a circle. Third, we quantify why this solution prevails over NC: a simplex ETF gains only an $O(1)$ advantage in cross-entropy, whereas the cyclic rank-2 solution enjoys a $Θ(K)$ advantage under Schatten or weight-decay surrogates, yielding a critical threshold $λ_{\mathrm{crit}} = Θ(1/K)$. Our results explain both why classifier weights move first and why embeddings subsequently align with them, showing that grokking on modular arithmetic is governed not by maximal separation alone but by a task-structured trade-off between separation, symmetry, and complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。