arXiv:2504.11713cs.LGcs.AI2025-04ICML被引 76

提出可高效扩展的扩散采样方法,显著提升采样规模与效率。

Adjoint Sampling: Highly Scalable Diffusion Samplers via Adjoint Matching

  • 基于伴随匹配原理,实现更多梯度更新而无需增加能量评估次数。
  • 在分子构象生成中成功应用,支持笛卡尔与扭转坐标系建模。
  • 适合需要大规模高效采样的计算化学研究者使用。

我们提出 Adjoint Sampling,一种高效且高度可扩展的扩散采样算法,用于从非归一化密度或能量函数中采样。它是首个在策略方法中允许梯度更新次数远超能量评估和模型采样次数的方法,使问题规模可显著扩展至以往方法难以触及的范围。该框架建立在随机最优控制理论基础上,具备与伴随匹配相同的理论保证,训练过程无需引入将样本推向目标分布的修正措施。我们展示了如何在笛卡尔和扭转坐标系中融入关键对称性及周期性边界条件,以建模分子系统。通过在经典能量函数上的广泛实验验证了方法的有效性,并进一步扩展至基于神经网络的能量模型,在多个分子体系上实现了高效的构象生成。为推动高可扩展采样方法的研究,我们将开源这些具有挑战性的基准测试,成功方法可直接促进计算化学的发展。

原文摘要 · Abstract (English)

We introduce Adjoint Sampling, a highly scalable and efficient algorithm for learning diffusion processes that sample from unnormalized densities, or energy functions. It is the first on-policy approach that allows significantly more gradient updates than the number of energy evaluations and model samples, allowing us to scale to much larger problem settings than previously explored by similar methods. Our framework is theoretically grounded in stochastic optimal control and shares the same theoretical guarantees as Adjoint Matching, being able to train without the need for corrective measures that push samples towards the target distribution. We show how to incorporate key symmetries, as well as periodic boundary conditions, for modeling molecules in both cartesian and torsional coordinates. We demonstrate the effectiveness of our approach through extensive experiments on classical energy functions, and further scale up to neural network-based energy models where we perform amortized conformer generation across many molecular systems. To encourage further research in developing highly scalable sampling methods, we plan to open source these challenging benchmarks, where successful methods can directly impact progress in computational chemistry.

扩散模型分子生成采样效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。