构建对称性分析工具,揭示神经网络优化中一阶与二阶结构的关联规律。
An Equivariance Toolbox for Learning Dynamics
- 提出统一框架,将梯度与海森矩阵约束结合,扩展对称性分析到离散变换
- 揭示损失曲面的平坦或陡峭方向与变换结构的对应关系,预测优化轨迹几何
- 适用于理解现代优化现象,适合研究模型泛化与训练动力学的学者
深度学习中的许多理论结果可追溯至神经网络在参数变换下的对称性或等变性。然而,现有分析通常局限于特定问题,且仅关注一阶后果(如守恒律),对二阶结构的影响理解不足。本文构建了一个通用的等变性工具箱,能推导出学习动态中耦合的一阶与二阶约束。该框架在三个方向扩展了经典诺特定理分析:从梯度约束扩展到海森约束,从对称性扩展到一般等变性,从连续变换扩展到离散变换。在一阶层面,框架将守恒律与隐式偏差关系统一为同一恒等式的特例;在二阶层面,它对曲率提供了结构性预测:哪些方向平坦或陡峭,梯度如何对齐海森特征空间,以及损失景观几何如何反映底层变换结构。通过多个应用展示了该框架的有效性,既复现了已知结果,也推导出新表征,连接变换结构与现代优化几何的实证观察。
原文摘要 · Abstract (English)
Many theoretical results in deep learning can be traced to symmetry or equivariance of neural networks under parameter transformations. However, existing analyses are typically problem-specific and focus on first-order consequences such as conservation laws, while the implications for second-order structure remain less understood. We develop a general equivariance toolbox that yields coupled first- and second-order constraints on learning dynamics. The framework extends classical Noether-type analyses in three directions: from gradient constraints to Hessian constraints, from symmetry to general equivariance, and from continuous to discrete transformations. At the first order, our framework unifies conservation laws and implicit-bias relations as special cases of a single identity. At the second order, it provides structural predictions about curvature: which directions are flat or sharp, how the gradient aligns with Hessian eigenspaces, and how the loss landscape geometry reflects the underlying transformation structure. We illustrate the framework through several applications, recovering known results while also deriving new characterizations that connect transformation structure to modern empirical observations about optimization geometry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。