用自洽的值函数指导扩散模型,提升生成质量与对齐效果
Explicit Critic Guidance for Aligning Diffusion Models

- 让扩散模型自身充当分时值函数,直接预测噪声状态价值
- 在单/多奖励任务中优于已有强化学习方法,测试时可调优生成质量
- 适合需要精细控制生成过程的研究者,尤其适用于非可微目标对齐
在线强化学习正日益成为对齐扩散模型与不可微目标的重要手段。然而,现有方法在去噪轨迹上的细粒度信用分配和基于值函数的稳定优化方面仍存在局限。我们提出一种状态对齐的潜在演员-评论家框架用于扩散模型后训练,其中扩散模型作为自身的分时条件值函数,直接在噪声潜在状态上预测价值。该设计支持轨迹级PPO训练,通过简单条件化与价值预训练策略实现稳定的演员-评论家优化,并自然支持学习到的评论家在推理时用于引导生成。进一步扩展至多奖励优化,联合训练互补奖励有助于缓解奖励黑客问题。在UNet与DiT双类骨干网络上,本方法在单奖励与多奖励基准上均持续优于先前的组相对强化学习及演员-评论家基线,且推理时的引导可进一步提升生成质量。
原文摘要 · Abstract (English)
Online reinforcement learning is becoming increasingly important for aligning diffusion models with non-differentiable objectives. However, existing methods still face limitations in assigning fine-grained credit along denoising trajectories and in realizing stable value-based optimization. We propose a state-aligned latent actor-critic framework for diffusion post-training, in which the diffusion model serves as its own timestep-conditioned value function and predicts values directly on noisy latent states. This enables trajectory-level PPO training, supports stable actor-critic optimization with simple conditioning and value pretraining strategies, and naturally allows the learned critic to be reused for inference-time steering. We further extend the framework to multi-reward optimization, where joint training with complementary rewards helps alleviate reward hacking. Across both UNet- and DiT-based backbones, our method consistently outperforms prior group-relative RL and actor-critic baselines on single-reward and multi-reward benchmarks, while test-time steering provides additional gains in generation quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。