构建标准化分子动力学评估框架,提升模型对比的可靠性与可重复性。
A Standardized Benchmark for Machine-Learned Molecular Dynamics using Weighted Ensemble Sampling
- 基于加权系综采样与TICA进展坐标,高效探索蛋白质构象空间。
- 涵盖9种不同蛋白、每种100万步模拟(4纳秒),支持19+项指标评估。
- 开源平台兼容经典力场与机器学习模型,适合算法开发者与验证者使用。
分子动力学(MD)方法的快速发展,尤其是机器学习动力学,已超过标准化评估工具的演进速度。方法间的客观比较常受制于评价指标不一致、罕见构象状态采样不足,以及缺乏可复现的基准测试。为解决这些问题,我们提出一个模块化基准测试框架,利用基于时间滞后独立分量分析(TICA)的进展坐标,通过并行化与分析工具(WESTPA)实现加权系综(WE)采样,系统评估蛋白质分子动力学方法。该框架包含轻量级可扩展的传播器接口,支持任意模拟引擎,兼容经典力场与机器学习模型。同时提供包含19项以上指标与可视化功能的全面评估套件。我们还贡献了9个多样蛋白数据集,长度从10到224个残基,覆盖不同折叠复杂度与拓扑结构。每个蛋白在300K下以每起始点100万步(4纳秒)进行充分模拟。通过使用隐式溶剂经典模拟作为对照,验证了全训练与欠训练的CGSchNet模型在构象采样上的差异。该开源平台通过统一评估协议,为分子模拟领域建立一致、严谨的基准测试基础。
原文摘要 · Abstract (English)
The rapid evolution of molecular dynamics (MD) methods, including machine-learned dynamics, has outpaced the development of standardized tools for method validation. Objective comparison between simulation approaches is often hindered by inconsistent evaluation metrics, insufficient sampling of rare conformational states, and the absence of reproducible benchmarks. To address these challenges, we introduce a modular benchmarking framework that systematically evaluates protein MD methods using enhanced sampling analysis. Our approach uses weighted ensemble (WE) sampling via The Weighted Ensemble Simulation Toolkit with Parallelization and Analysis (WESTPA), based on progress coordinates derived from Time-lagged Independent Component Analysis (TICA), enabling fast and efficient exploration of protein conformational space. The framework includes a flexible, lightweight propagator interface that supports arbitrary simulation engines, allowing both classical force fields and machine learning-based models. Additionally, the framework offers a comprehensive evaluation suite capable of computing more than 19 different metrics and visualizations across a variety of domains. We further contribute a dataset of nine diverse proteins, ranging from 10 to 224 residues, that span a variety of folding complexities and topologies. Each protein has been extensively simulated at 300K for one million MD steps per starting point (4 ns). To demonstrate the utility of our framework, we perform validation tests using classic MD simulations with implicit solvent and compare protein conformational sampling using a fully trained versus under-trained CGSchNet model. By standardizing evaluation protocols and enabling direct, reproducible comparisons across MD approaches, our open-source platform lays the groundwork for consistent, rigorous benchmarking across the molecular simulation community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。