用增强采样提升粗粒化机器学习势的训练效率与精度
Enhanced Sampling for Efficient Learning of Coarse-Grained Machine Learning Potentials
- 通过沿粗粒化自由度施加偏置,加速生成平衡数据
- 在不改变势能平均力的前提下,显著改善过渡区采样
- 适合追求高精度粗粒化模拟的研究者使用
粗粒化(CG)使分子动力学模拟得以处理更大系统和更长时间尺度,而传统原子模型难以实现。机器学习势(MLPs)可捕捉多体相互作用,为粗粒化模型中的平均力势(PMF)提供精确近似。当前的CG MLP通常采用自下而上的力匹配方法训练,依赖无偏平衡玻尔兹曼分布采样以保证热力学一致性。这一做法存在两大局限:一是需足够长的原子轨迹才能收敛;二是即使达到平衡,过渡区域仍采样不足。为此,本文采用增强采样技术,在粗粒化自由度上施加偏置以生成数据,再基于无偏势能重新计算力。该策略同时缩短了生成平衡数据所需时间,并丰富了过渡区域的采样,且保持正确的PMF。我们在Müller-Brown势和封端丙氨酸体系上验证了其有效性,取得了显著改进。结果表明,将增强采样用于力匹配是提升CG MLP准确性和可靠性的有前景方向。
原文摘要 · Abstract (English)
Coarse-graining (CG) enables molecular dynamics (MD) simulations of larger systems and longer timescales that are otherwise infeasible with atomistic models. Machine learning potentials (MLPs), with their capacity to capture many-body interactions, can provide accurate approximations of the potential of mean force (PMF) in CG models. Current CG MLPs are typically trained in a bottom-up manner via force matching, which in practice relies on configurations sampled from the unbiased equilibrium Boltzmann distribution to ensure thermodynamic consistency. This convention poses two key limitations: first, sufficiently long atomistic trajectories are needed to reach convergence; and second, even once equilibrated, transition regions remain poorly sampled. To address these issues, we employ enhanced sampling to bias along CG degrees of freedom for data generation, and then recompute the forces with respect to the unbiased potential. This strategy simultaneously shortens the simulation time required to produce equilibrated data and enriches sampling in transition regions, while preserving the correct PMF. We demonstrate its effectiveness on the Müller-Brown potential and capped alanine, achieving notable improvements. Our findings support the use of enhanced sampling for force matching as a promising direction to improve the accuracy and reliability of CG MLPs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。