arXiv:2606.01636cs.CV2026-06

通过速度分解提升生成模型的偏好对齐精度

Pave-GRPO: Beyond Instantaneous Guidance through Principled Average Velocity Decomposition

论文配图:Pave-GRPO: Beyond Instantaneous Guidance through Principled Average Velocity Decomposition
图 1 · 摘自论文原文
  • 用解析式分解粗轨迹为细子路径,实现高效密集监督
  • 仅需少量额外计算,就能扩展优化范围并覆盖更多中间阶段
  • 适合追求高质量生成与高效训练的扩散模型研究者

组相对策略优化(GRPO)已成为对齐基于流的生成模型与人类偏好的有效范式。然而,群体回放的高成本迫使现有方法采用极少的去噪步数,导致时间监督稀疏,大部分中间阶段缺乏直接奖励引导。为此,我们提出Pave-GRPO,通过原理性的平均速度分解重构GRPO目标。无需生成昂贵的高步数回放,我们维持高效的少步群体采样,但将每个粗略转移分解为跨越多个中间时间步的等效细粒度子轨迹集合,将奖励反馈传播至更密集的时间阶段,实现更全面的偏好对齐。关键在于,此过程不增加随机回放生成或奖励评估开销:分解后的子轨迹在观测转移周围解析构造,仅需在策略更新时进行少量额外的速度网络评估。该设计带来双重优势:(i) 无回放的时域扩展:通过直接复用少步群体样本及其关联奖励,Pave-GRPO在固定采样与奖励预算下显著扩大有效优化范围;(ii) 全面的时间监督:通过将瞬时速度目标等效分解为多时间步集合,使奖励信号分布于去噪过程的更多中间阶段,实现更精细、更彻底的偏好优化。

原文摘要 · Abstract (English)

Group Relative Policy Optimization(GRPO) has emerged as an effective paradigm for aligning flow-based generative models with human preferences. However, the high cost of group rollouts forces existing methods to use very few denoising steps, resulting in sparse temporal supervision and leaving most intermediate stages without direct reward guidance. To address this, we propose Pave-GRPO, which reformulates the GRPO objective through principled average velocity decomposition. Rather than generating expensive high-step rollouts, we maintain efficient few-step group sampling but decompose each coarse transition into an equivalent ensemble of finer sub-trajectories spanning multiple intermediate timesteps, propagating reward feedback to a denser set of temporal stages for more comprehensive preference alignment. Crucially, this incurs no additional stochastic rollout generation or reward evaluation: the decomposed sub-trajectories are constructed analytically around the observed transitions, requiring only a small number of extra velocity-network evaluations during the policy update. This design offers two benefits: (i) rollout-free horizon expansion: through the direct reuse of few-step group samples and their associated rewards, Pave-GRPO significantly broadens the effective optimization scope under a fixed sampling and reward budget; and (ii) comprehensive temporal supervision: by equivalently decomposing an instantaneous velocity target into a multi-timestep ensemble, it distributes reward signals across more intermediate stages of the denoising process, enabling finer-grained and more thorough preference optimization.

生成模型偏好对齐扩散模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。