用方差最小化优化扩散模型采样,更稳定高效。
Diffusion Alignment Beyond KL: Variance Minimisation as Effective Policy Optimiser
- 以重要性权重方差最小化替代传统KL目标
- 理论证明方差目标在最优时对应奖励倾斜分布
- 统一解释多种方法并启发新设计方向
扩散对齐将预训练扩散模型调整为沿去噪轨迹从奖励倾斜分布中采样。该过程自然可视为顺序蒙特卡洛(SMC),其中去噪模型作为提议分布,奖励引导产生重要性权重。受此视角启发,我们提出方差最小化策略优化(VMPO),将扩散对齐表述为最小化对数重要性权重的方差,而非直接优化基于KL的损失。我们证明,方差目标在奖励倾斜目标分布处取得最小值,且在同策略采样下,其梯度与标准KL对齐一致。这一视角为理解扩散对齐提供了统一框架。通过选择不同的势函数和方差最小化策略,VMPO可恢复多种现有方法,并揭示超越KL的新设计方向。
原文摘要 · Abstract (English)
Diffusion alignment adapts pretrained diffusion models to sample from reward-tilted distributions along the denoising trajectory. This process naturally admits a Sequential Monte Carlo (SMC) interpretation, where the denoising model acts as a proposal and reward guidance induces importance weights. Motivated by this view, we introduce Variance Minimisation Policy Optimisation (VMPO), which formulates diffusion alignment as minimising the variance of log importance weights rather than directly optimising a Kullback-Leibler (KL) based objective. We prove that the variance objective is minimised by the reward-tilted target distribution and that, under on-policy sampling, its gradient coincides with that of standard KL-based alignment. This perspective offers a common lens for understanding diffusion alignment. Under different choices of potential functions and variance minimisation strategies, VMPO recovers various existing methods, while also suggesting new design directions beyond KL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。