arXiv:2602.19533cs.LGcs.AI2026-02被引 1

研究神经网络在代数运算中的突然泛化现象,揭示数学结构如何决定学习过程。

Grokking Finite-Dimensional Algebra

  • 将神经网络的突现泛化现象扩展到非交换、非结合等更一般代数结构。
  • 发现实数域上学习对应矩阵低秩分解,有限域上自然产生离散表示学习。
  • 适合对神经网络泛化机制与数学结构关系感兴趣的学者。

本文研究神经网络训练中从长期记忆到突然泛化的突现现象(grokking),聚焦于有限维代数(FDA)中的乘法学习。以往研究集中于群运算,本文将其拓展至非结合、非交换、无单位元的代数结构。我们证明群运算是学习FDA的特例,乘法学习本质是学习由代数结构张量定义的双线性映射。对于实数域上的代数,该学习问题与具有隐式低秩偏置的矩阵分解相关;对于有限域上的代数,突现泛化自然出现,因模型需学习代数元素的离散表示。我们实验探究了三个核心问题:(i) 交换性、结合性、单位元等代数性质如何影响突现泛化的出现时机;(ii) 结构张量的稀疏性和秩如何影响泛化能力;(iii) 泛化程度是否与模型学习到与代数表示对齐的潜在嵌入相关。本工作为不同代数结构下的突现泛化提供了统一框架,并揭示数学结构如何调控神经网络泛化动态。

原文摘要 · Abstract (English)

This paper investigates the grokking phenomenon, which refers to the sudden transition from a long memorization to generalization observed during neural networks training, in the context of learning multiplication in finite-dimensional algebras (FDA). While prior work on grokking has focused mainly on group operations, we extend the analysis to more general algebraic structures, including non-associative, non-commutative, and non-unital algebras. We show that learning group operations is a special case of learning FDA, and that learning multiplication in FDA amounts to learning a bilinear product specified by the algebra's structure tensor. For algebras over the reals, we connect the learning problem to matrix factorization with an implicit low-rank bias, and for algebras over finite fields, we show that grokking emerges naturally as models must learn discrete representations of algebraic elements. This leads us to experimentally investigate the following core questions: (i) how do algebraic properties such as commutativity, associativity, and unitality influence both the emergence and timing of grokking, (ii) how structural properties of the structure tensor of the FDA, such as sparsity and rank, influence generalization, and (iii) to what extent generalization correlates with the model learning latent embeddings aligned with the algebra's representation. Our work provides a unified framework for grokking across algebraic structures and new insights into how mathematical structure governs neural network generalization dynamics.

神经网络代数结构泛化机制突现学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。