新优化器让机器学习势函数训练更快更准,尤其适合标注数据少的情况。
Beyond Adam: SOAP and Muon for Faster, Label-Efficient Training of Machine Learning Interatomic Potentials

- 用矩阵结构优化器替代Adam,提升训练效率
- 在部分力监督下,收敛速度和精度显著优于Adam
- SOAP与混合方法表现稳定,适合科研高效建模
机器学习互作用势(MLIPs)已成为人工智能在科学模拟中的标志。尽管架构与数据集的改进已使模型精度和泛化能力持续提升,但训练优化器的选择仍基本未受关注,社区普遍沿用Adam及其变体。本文系统实现并比较了近期提出的几类矩阵结构优化器——Muon、SOAP及混合型SOAP-Muon,用于训练NequIP与Allegro MLIP模型。结果表明,这些优化器在收敛速度和最终精度上均显著优于Adam。其中,SOAP与SOAP-Muon表现稳健且持续领先,而Muon仅带来部分提升。上述优势在部分力监督条件下尤为突出。研究揭示,优化器选择是MLIP设计中被忽视却极具影响力的维度。
原文摘要 · Abstract (English)
Machine learning interatomic potentials (MLIPs) have become a hallmark of AI for scientific simulation. While efforts on new architectures and datasets have led to increasingly accurate and general models, the choice of optimizer for training has largely remained unexplored, defaulting to Adam and its variants in the community. Here, we implement and systematically compare a class of recently proposed matrix-structured optimizers, including Muon, SOAP, and the hybrid SOAP-Muon, for training NequIP and Allegro MLIP models. We find that these optimizers can substantially outperform Adam in both convergence speed and final accuracy. SOAP and SOAP-Muon emerge as robust and consistently strong methods, while Muon only provides partial gains relative to Adam. The improvements are particularly pronounced under partial force supervision. Our results indicate that optimizer choice is an overlooked yet impactful design axis for MLIPs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。