无需数据即可生成分子模拟的精细结构,直接逼近真实能量分布。
Energy-Based Coarse-Graining in Molecular Dynamics: A Flow-Based Framework without Data
- 基于能量目标构建无数据生成框架,直接建模原子级玻尔兹曼分布。
- 在双阱和丙氨酸二肽系统中成功捕获所有能态,实现原子级结构重建。
- 适合需要高精度反向映射的分子模拟研究,尤其适用于缺乏训练数据场景。
粗粒化(CG)模型可有效降低分子动力学模拟的复杂度,但传统方法严重依赖长时间全原子轨迹以充分采样构型空间,导致模型准确性和泛化能力受限,未访问的构型无法被包含。本文提出一种完全无数据的生成式粗粒化框架,直接针对全原子玻尔兹曼分布进行建模。该模型定义了一个结构化潜空间,包含慢变集体变量(对应多模态边缘分布,捕捉亚稳态)和快变变量(通过简单单峰条件分布表示)。一个可学习的双射映射从潜空间到原子坐标,实现分子结构的自动精确重构。训练仅依赖原子间势能,通过最小化反向KL散度优化,采用自适应退火策略稳定优化并促进多样构型探索。模型训练完成后,可一次性生成独立的平衡态全原子样本。在双阱势和高斯混合模型两个合成系统以及标准测试体系丙氨酸二肽上的验证表明,该方法能准确捕获玻尔兹曼分布的所有相关模式,重建原子构型,并自动学习物理上合理的粗粒化表示。结果表明,该方法为传统粗粒化技术提供了一种有原则的数据自由替代方案,同时解决了长期存在的‘鸡与蛋’难题,并通过精准重构全原子构型有效解决了反向映射问题。
原文摘要 · Abstract (English)
Coarse-grained (CG) models provide an effective route to reducing the complexity of molecular simulations (MD), but conventional approaches depend heavily on long all-atom MD trajectories to adequately sample configurational space. This data dependence limits accuracy and generalizability, as unvisited configurations remain excluded from the resulting CG models. We introduce a fully data-free, generative framework for CG that directly targets the all-atom Boltzmann distribution. The model defines a structured latent space comprising slow collective variables, associated with multimodal marginal densities capturing metastable states, and fast variables, represented through simple, unimodal conditional distributions. A learnable, bijective map from latent space to atomistic coordinates enables the automatic and accurate reconstruction of molecular structures. Training relies solely on the interatomic potential and minimizes the reverse Kullback-Leibler (KL) divergence via an energy-based objective. To stabilize optimization and ensure mode coverage, we employ an adaptive tempering scheme that promotes the exploration of diverse configurations. Once trained, the model can generate independent, one-shot equilibrium samples at full atomic resolution. Validation on two synthetic systems, a double-well potential and a Gaussian mixture model, as well as on the benchmark alanine dipeptide, demonstrates that the method captures all relevant modes of the Boltzmann distribution, reconstructs atomic configurations, and automatically learns physically meaningful CG representations. These results suggest a promising, data-free alternative to traditional CG techniques, offering both a principled approach to addressing the long-standing "chicken-and-egg" challenge in coarse-graining and an effective solution to the back-mapping problem by enabling accurate reconstruction of all-atom configurations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。