arXiv:2508.18175cs.LGcs.AI2025-08NeurIPS被引 17

用深度学习构建可迁移的分子构象采样器,一次训练通用所有肽链。

Amortized Sampling with Transferable Normalizing Flows

  • 基于肽链轨迹训练2.85亿参数归一化流模型,实现零样本跨序列采样。
  • 在任意肽链上生成不相关采样样本,支持长度扩展且保持高效似然计算。
  • 适用于多种采样算法,开源代码与数据,推动自适应采样研究。

高效的分子构象平衡采样仍是计算化学与统计推断的核心挑战。传统方法如分子动力学或马尔可夫链蒙特卡洛缺乏自适应性,每个系统需独立耗时采样。生成模型的成功激发了学习采样算法的探索。尽管在单一系统上表现媲美传统方法,现有学习采样器跨系统迁移能力有限。本文提出Prose——一个2.85亿参数的全原子可迁移归一化流模型,基于最长8个残基的肽链分子动力学轨迹训练。Prose可在任意肽链上实现零样本、不相关提议采样,首次实现序列长度上的迁移性,同时保留归一化流的高效似然评估能力。通过广泛实验证明,Prose作为多种采样算法的提议分布有效,简单的基于重要性采样的微调即可达到顺序蒙特卡洛等成熟方法的竞争力。我们开源了Prose的代码库、模型权重与训练数据集,以进一步推动自适应采样方法与目标的研究。

原文摘要 · Abstract (English)

Efficient equilibrium sampling of molecular conformations remains a core challenge in computational chemistry and statistical inference. Classical approaches such as molecular dynamics or Markov chain Monte Carlo inherently lack amortization; the computational cost of sampling must be paid in full for each system of interest. The widespread success of generative models has inspired interest towards overcoming this limitation through learning sampling algorithms. Despite performing competitively with conventional methods when trained on a single system, learned samplers have so far demonstrated limited ability to transfer across systems. We demonstrate that deep learning enables the design of scalable and transferable samplers by introducing Prose, a 285 million parameter all-atom transferable normalizing flow trained on a corpus of peptide molecular dynamics trajectories up to 8 residues in length. Prose draws zero-shot uncorrelated proposal samples for arbitrary peptide systems, achieving the previously intractable transferability across sequence length, whilst retaining the efficient likelihood evaluation of normalizing flows. Through extensive empirical evaluation we demonstrate the efficacy of Prose as a proposal for a variety of sampling algorithms, finding a simple importance sampling-based fine-tuning procedure to achieve competitive performance to established methods such as sequential Monte Carlo. We open-source the Prose codebase, model weights, and training dataset, to further stimulate research into amortized sampling methods and objectives.

分子采样归一化流可迁移生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。