用可微分代理奖励提升两步扩散模型的对齐效果
Reward Fine-Tuning Two-Step Diffusion Models via Learning Differentiable Latent-Space Surrogate Reward
- 在潜在空间学习可微代理奖励,将任意奖励转为可导形式
- 在≤2步生成中显著优于DDPO、Diffusion-DPO等方法
- 适合追求超快生成且需精确对齐的应用场景
近期研究证明,通过强化学习(RL)技术可对扩散模型(DMs)进行任意奖励的微调,包括不可导奖励,实现灵活对齐。然而,现有RL方法在超快速(≤2步)图像生成的步骤压缩扩散模型上应用困难。分析表明,基于策略的RL方法(如PPO或DPO)存在局限性。为此,我们提出LaSRO:在SDXL的潜在空间中学习可微代理奖励模型,将任意奖励转换为可导形式,实现有效的梯度引导。该方法利用预训练潜在扩散模型进行奖励建模,并针对≤2步生成优化奖励,支持高效非策略探索。实验表明,LaSRO在多种奖励目标下均有效且稳定,显著优于主流的DDPO与Diffusion-DPO。此外,我们揭示了其与基于价值的强化学习的联系,提供理论支持。
原文摘要 · Abstract (English)
Recent research has shown that fine-tuning diffusion models (DMs) with arbitrary rewards, including non-differentiable ones, is feasible with reinforcement learning (RL) techniques, enabling flexible model alignment. However, applying existing RL methods to step-distilled DMs is challenging for ultra-fast ($\le2$-step) image generation. Our analysis suggests several limitations of policy-based RL methods such as PPO or DPO toward this goal. Based on the insights, we propose fine-tuning DMs with learned differentiable surrogate rewards. Our method, named LaSRO, learns surrogate reward models in the latent space of SDXL to convert arbitrary rewards into differentiable ones for effective reward gradient guidance. LaSRO leverages pre-trained latent DMs for reward modeling and tailors reward optimization for $\le2$-step image generation with efficient off-policy exploration. LaSRO is effective and stable for improving ultra-fast image generation with different reward objectives, outperforming popular RL methods including DDPO and Diffusion-DPO. We further show LaSRO's connection to value-based RL, providing theoretical insights. See our webpage \href{https://sites.google.com/view/lasro}{here}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。