提出新型非局部力场模型,让大分子模拟更准更快。
Scalable Machine Learning Force Fields for Macromolecular Systems Through Long-Range Aware Message Passing
- 用显式长程注意力机制突破传统模型局部限制
- 在1200原子系统上误差不随规模增长,精度显著提升
- 适合蛋白质等大分子体系的高保真模拟,效率更高
机器学习力场(MLFF)通过量子力学精度实现分子模拟的快速计算,但其依赖固定截断范围的架构限制了在以长程相互作用为主的生物大分子系统中的应用。我们发现,这种局域性导致力预测误差随体系增大而单调上升,暴露出关键的架构瓶颈。为此,构建了包含至1200个原子的高保真密度泛函理论(DFT)基准数据集MolLR25,提出E2Former-LSR——一种具备对称性保持特性的变换器模型,显式引入长程注意力模块。该模型实现误差稳定增长,精准捕捉非共价相互作用衰减,并在复杂蛋白构象上保持高精度。更重要的是,其高效设计相较纯局部模型提升约30%运算速度。本工作验证了非局部架构对通用化机器学习力场的必要性,推动了大规模化学与生物体系的高保真分子动力学模拟。
原文摘要 · Abstract (English)
Machine learning force fields (MLFFs) have revolutionized molecular simulations by providing quantum mechanical accuracy at the speed of molecular mechanical computations. However, a fundamental reliance of these models on fixed-cutoff architectures limits their applicability to macromolecular systems where long-range interactions dominate. We demonstrate that this locality constraint causes force prediction errors to scale monotonically with system size, revealing a critical architectural bottleneck. To overcome this, we establish the systematically designed MolLR25 ({Mol}ecules with {L}ong-{R}ange effect) benchmark up to 1200 atoms, generated using high-fidelity DFT, and introduce E2Former-LSR, an equivariant transformer that explicitly integrates long-range attention blocks. E2Former-LSR exhibits stable error scaling, achieves superior fidelity in capturing non-covalent decay, and maintains precision on complex protein conformations. Crucially, its efficient design provides up to 30% speedup compared to purely local models. This work validates the necessity of non-local architectures for generalizable MLFFs, enabling high-fidelity molecular dynamics for large-scale chemical and biological systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。