优化器影响等变神经网络的训练效果,Muon比Adam更优。
How the Optimizer Shapes Learned Solutions in Equivariant Neural Networks
- 对比Adam与Muon在等变网络中的表现
- Muon使模型损失曲面更规则,权重秩更高
- 适合研究几何归纳偏置与优化器关系的学者
等变神经网络通过构造编码几何对称性,但通常难以优化且性能低于无约束架构。现有工作多从网络结构入手,如放宽约束或近似等变性,而优化器的作用仍被忽视。本文在点云和分子学习任务中,对比了Muon与Adam在多种等变及几何架构上的表现。在ModelNet40上,Muon在所有架构下均优于Adam。进一步分析表明,Muon训练出的模型具有更高的海森矩阵曲率总和、更规则的损失曲面,且学习到的权重与中间表示具有更高的稳定秩与有效秩。这些结果表明,优化器设计与几何归纳偏置之间的相互作用值得更多关注。
原文摘要 · Abstract (English)
Equivariant neural networks encode geometric symmetries by construction, yet they are often difficult to optimize and can underperform less constrained architectures. A growing body of work addresses this through architectural modifications such as constraint relaxation or approximate equivariance, while the role of the optimizer remains comparatively underexplored. We study this direction by comparing Muon and Adam across several equivariant and geometric architectures under pointcloud and molecular learning settings. On ModelNet40, where the comparison is clearest, Muon consistently improves over Adam across all architectures considered. We then analyze the trained ModelNet40 checkpoints through Hessian estimates, loss surface visualizations, and spectral properties of learned weights and intermediate representations. The checkpoints reached by Muon have larger Hessian curvature summaries but more regular loss surfaces, and their learned weights and representations have higher stable and effective ranks. These observations suggest that the interaction between optimizer design and geometric inductive bias deserves further attention from the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。