用新指标检测机器学习势能面不光滑问题,提升模型设计效率。
From Evaluation to Design: Using Potential Energy Surface Smoothness Metrics to Guide Machine Learning Interatomic Potential Architectures
- 提出基于键变形的高效评测方法BSCT,探测势能面非平滑性。
- BSCT与分子动力学稳定性强相关,计算成本仅为后者的1/10。
- 可指导模型迭代优化,适合开发高性能机器学习势的科研人员。
机器学习原子间势(MLIPs)有时无法复现量子势能面(PES)的物理平滑性,导致下游模拟中出现错误行为,而传统能量和力回归评估难以发现此类问题。现有评估方法如微正则分子动力学(MD)计算成本高,且主要探测平衡附近状态。为改进评估指标,本文提出键平滑性表征测试(BSCT),通过受控键变形探测远近平衡态下的势能面非平滑性,包括不连续、人工极小值和异常力。实验表明,BSCT与MD稳定性高度相关,但计算开销仅为后者的约十分之一。以无约束Transformer为测试平台,我们展示了通过引入新的可微分k-近邻算法和温度控制注意力机制,可有效降低BSCT识别出的伪影。基于BSCT系统优化模型设计后,所获MLIP在常规能量/力回归误差低的同时,实现稳定分子动力学模拟与鲁棒原子性质预测。结果表明,BSCT既可作为实践者评估MLIP实用性的验证工具,也可作为开发者在模型迭代中实时反馈物理挑战的“闭环”设计代理。代码与数据集已开源:https://github.com/ryanliu30/bsct.git。
原文摘要 · Abstract (English)
Machine Learning Interatomic Potentials (MLIPs) sometimes fail to reproduce the physical smoothness of the quantum potential energy surface (PES), leading to erroneous behavior in downstream simulations that standard energy and force regression evaluations can miss. Existing evaluations, such as microcanonical molecular dynamics (MD), are computationally expensive and primarily probe near-equilibrium states. To improve evaluation metrics for MLIPs, we introduce the Bond Smoothness Characterization Test (BSCT). This efficient benchmark probes the PES via controlled bond deformations and detects non-smoothness, including discontinuities, artificial minima, and spurious forces, both near and far from equilibrium. We show that BSCT correlates strongly with MD stability while requiring a fraction of the cost of MD. To demonstrate how BSCT can guide iterative model design, we utilize an unconstrained Transformer backbone as a testbed, illustrating how refinements such as a new differentiable $k$-nearest neighbors algorithm and temperature-controlled attention reduce artifacts identified by our metric. By optimizing model design systematically based on BSCT, the resulting MLIP simultaneously achieves a low conventional E/F regression error, stable MD simulations, and robust atomistic property predictions. Our results establish BSCT as both a validation metric for practitioners to assess MLIP utility and as an "in-the-loop" model design proxy that alerts MLIP developers to physical challenges that cannot be efficiently evaluated by current MLIP benchmarks. The BSCT dataset and evaluation are available on https://github.com/ryanliu30/bsct.git
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。