长程相互作用建模能显著提升机器学习势函数的泛化能力。
Generalization of Long-Range Machine Learning Potentials in Complex Chemical Spaces
- 引入长程校正机制,增强模型对复杂化学空间的适应性。
- 在金属有机框架上实现跨分布样本的性能提升,验证了迁移能力突破。
- 适合材料模拟、分子设计等领域研究者参考模型设计策略。
化学空间的庞大使得泛化成为机器学习原子间势(MLIPs)发展的核心挑战。尽管MLIPs可实现接近量子精度的大规模原子模拟,但其应用常受限于对分布外样本的差强人意的迁移能力。本文系统评估了多种带长程校正的MLIP架构在多样化化学空间中的表现,结果表明此类方案不仅提升分布内性能,更关键的是显著增强对未见化学区域的泛化能力。为实现更严格的基准测试,我们提出有偏训练-测试划分策略,明确检验模型在显著不同的化学空间区域的表现。综合成果强调了长程建模在实现可泛化MLIPs中的重要性,并提供诊断化学空间系统性失败的框架。尽管我们在金属有机框架上验证方法,但其适用范围广泛,为构建更鲁棒、可迁移的MLIPs提供了洞见。
原文摘要 · Abstract (English)
The vastness of chemical space makes generalization a central challenge in the development of machine learning interatomic potentials (MLIPs). While MLIPs could enable large-scale atomistic simulations with near-quantum accuracy, their usefulness is often limited by poor transferability to out-of-distribution samples. Here, we systematically evaluate different MLIP architectures with long-range corrections across diverse chemical spaces and show that such schemes are essential, not only for improving in-distribution performance but, more importantly, for enabling significant gains in transferability to unseen regions of chemical space. To enable a more rigorous benchmarking, we introduce biased train-test splitting strategies, which explicitly test the model performance in significantly different regions of chemical space. Together, our findings highlight the importance of long-range modeling for achieving generalizable MLIPs and provide a framework for diagnosing systematic failures across chemical space. Although we demonstrate our methodology on metal-organic frameworks, it is broadly applicable to other materials, offering insights into the design of more robust and transferable MLIPs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。