用价值函数梯度引导扩散采样,提升对称系统采样效率与质量
Value Gradient Sampler: Learning Invariant Value Functions for Equivariant Diffusion Sampling
- 用价值函数定义梯度流,实现对称性不变的采样路径
- 在55粒子伦纳德-琼斯系统上,采样速度最快、样本质量最高
- 兼容离线强化学习,适合复杂对称系统的高效采样
我们提出价值梯度采样器(VGS),一种由价值函数参数化的扩散采样方法。VGS通过沿价值函数梯度演化随机初始化的粒子,从非归一化目标密度(即能量)生成样本。在目标密度具有对称性不变性的许多采样问题中,价值函数提供了一种新思路:利用不变网络诱导等变梯度流,无需构建更复杂的等变网络。价值网络通过时序差分学习训练,支持离线策略训练及其他成熟强化学习技术。结合先进强化学习方法与高效不变网络,VGS在55粒子伦纳德-琼斯系统上,相较基线方法实现了最高的样本质量和最快的采样速度。
原文摘要 · Abstract (English)
We propose the Value Gradient Sampler (VGS), a diffusion sampler parameterized by value functions. VGS generates samples from an unnormalized target density (i.e., energy) by evolving randomly initialized particles along the gradient of the value function. In many sampling problems where the target density exhibits invariant symmetries, value functions provide a novel approach to leveraging invariant networks for sampling by inducing an equivariant gradient flow, without requiring more complex equivariant networks. The value networks are trained via temporal difference learning, which supports off-policy training and other established reinforcement learning (RL) techniques. By combining advanced RL methods with efficient invariant networks, VGS achieves both the highest sample quality and the fastest sampling speed among our baselines on the 55-particle Lennard-Jones system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。