发现算法任务内在对称性是模型从记忆走向泛化的关键驱动力。
Intrinsic Task Symmetry Drives Generalization in Algorithmic Tasks
- 通过识别记忆、对称性获取、几何组织三阶段,揭示泛化机制。
- 对称性获取阶段即出现泛化,随后表征空间形成任务对齐的结构。
- 适用于需要稳健算法推理的场景,如代数与关系推理任务。
Grokking 指的是模型从记忆到泛化的突然转变,其特征是低维表征的出现,但其组织机制仍不明确。我们提出,内在任务对称性主要驱动 grokking,并塑造模型表征空间的几何结构。我们识别出 grokking 背后的稳定三阶段训练动态:(i) 记忆阶段,(ii) 对称性获取阶段,(iii) 几何组织阶段。我们发现泛化在对称性获取阶段即已出现,此后表征重新组织为结构化、任务对齐的几何形态。我们在多种算法领域(包括代数、结构和关系推理任务)验证了这一对称性驱动的解释。基于此,我们提出了一个基于对称性的诊断方法,可提前预测泛化开始时间,并提出加速泛化的策略。结果表明,内在对称性是神经网络突破记忆、实现稳健算法推理的关键因素。
原文摘要 · Abstract (English)
Grokking, the sudden transition from memorization to generalization, is characterized by the emergence of low-dimensional representations, yet the mechanism underlying this organization remains elusive. We propose that intrinsic task symmetries primarily drive grokking and shape the geometry of the model's representation space. We identify a consistent three-stage training dynamic underlying grokking: (i) memorization, (ii) symmetry acquisition, and (iii) geometric organization. We show that generalization emerges during the symmetry acquisition phase, after which representations reorganize into a structured, task-aligned geometry. We validate this symmetry-driven account across diverse algorithmic domains, including algebraic, structural, and relational reasoning tasks. Building on these findings, we introduce a symmetry-based diagnostic that anticipates the onset of generalization and propose strategies to accelerate it. Together, our results establish intrinsic symmetry as the key factor enabling neural networks to move beyond memorization and achieve robust algorithmic reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。