跨功能数据迁移提升通用原子势能模型精度
Cross-functional transferability in universal machine learning interatomic potentials
- 在CHGNet框架下研究低到高保真度数据的迁移学习
- 元素能量参考可显著提升迁移效果,实现高效训练
- 适合关注多保真度学习与下一代势能模型的研究者
通用机器学习原子势能(uMLIPs)的快速发展展示了泛化学习通用势能面的可能性。理论上,通过将模型从低保真度数据迁移到高保真度数据,可进一步提升精度。本文在CHGNet框架下分析了这一迁移学习问题,发现GGA与r²SCAN之间存在显著的能量尺度偏移和弱相关性,阻碍了跨功能数据迁移。基于包含24万结构的MP-r²SCAN数据集,我们对比多种迁移学习方法,证实元素能量参考对uMLIP迁移学习至关重要。通过对比有无低保真度预训练的标度律,表明即使目标数据集不足百万结构,仍可通过迁移学习实现显著的数据效率提升。研究强调了恰当的迁移学习与多保真度学习在构建下一代高保真度uMLIP中的关键作用。
原文摘要 · Abstract (English)
The rapid development of universal machine learning interatomic potentials (uMLIPs) has demonstrated the possibility for generalizable learning of the universal potential energy surface. In principle, the accuracy of uMLIPs can be further improved by bridging the model from lower-fidelity datasets to high-fidelity ones. In this work, we analyze the challenge of this transfer learning problem within the CHGNet framework. We show that significant energy scale shifts and poor correlations between GGA and r$^2$SCAN pose challenges to cross-functional data transferability in uMLIPs. By benchmarking different transfer learning approaches on the MP-r$^2$SCAN dataset of 0.24 million structures, we demonstrate the importance of elemental energy referencing in the transfer learning of uMLIPs. By comparing the scaling law with and without the pre-training on a low-fidelity dataset, we show that significant data efficiency can still be achieved through transfer learning, even with a target dataset of sub-million structures. We highlight the importance of proper transfer learning and multi-fidelity learning in creating next-generation uMLIPs on high-fidelity data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。