用全局注意力机制突破长程作用力建模瓶颈,实现超大规模分子模拟
A recipe for scalable attention-based MLIPs: unlocking long-range accuracy with all-to-all node attention
- 采用全连接节点注意力机制,数据驱动捕捉长程相互作用
- 在1亿级样本上训练,能量/力精度达当前最优,且可稳定进行长时间分子动力学模拟
- 适合大体系分子系统模拟,尤其对生物分子和电解质等场景有显著优势
机器学习原子间势函数(MLIPs)发展迅速,许多顶尖模型依赖强物理先验。但当模型扩展至生物分子、电解质等大体系时,难以准确描述长程相互作用,现有方法多依赖显式物理项。本文提出AllScAIP,一种基于注意力机制、能量守恒的简洁MLIP模型,可扩展至约1亿样本训练。其通过全连接节点注意力组件实现数据驱动的长程作用力建模。大量消融实验表明,在小数据/小模型条件下,物理先验提升样本效率;但随数据与模型规模扩大,其优势减弱甚至逆转,而全连接注意力始终对捕捉长程相互作用至关重要。模型在分子系统上达到当前最优的能量与力精度,并在OMol25多项物理评估中表现优异,同时在材料(OMat24)和催化剂(OC20)任务上保持竞争力。此外,该模型支持稳定长时分子动力学模拟,能准确复现密度与汽化热等实验可观测量。
原文摘要 · Abstract (English)
Machine-learning interatomic potentials (MLIPs) have advanced rapidly, with many top models relying on strong physics-based inductive biases. However, as models scale to larger systems like biomolecules and electrolytes, they struggle to accurately capture long-range (LR) interactions, leading current approaches to rely on explicit physics-based terms or components. In this work, we propose AllScAIP, a straightforward, attention-based, and energy-conserving MLIP model that scales to O(100 million) training samples. It addresses the long-range challenge using an all-to-all node attention component that is data-driven. Extensive ablations reveal that in low-data/small-model regimes, inductive biases improve sample efficiency. However, as data and model size scale, these benefits diminish or even reverse, while all-to-all attention remains critical for capturing LR interactions. Our model achieves state-of-the-art energy/force accuracy on molecular systems, as well as a number of physics-based evaluations (OMol25), while being competitive on materials (OMat24) and catalysts (OC20). Furthermore, it enables stable, long-timescale MD simulations that accurately recover experimental observables, including density and heat of vaporization predictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。