分步赋信提升扩散模型强化学习效率
Stepwise Credit Assignment for GRPO on Flow-Matching Models
- 按每步的奖励改进分步赋信,更合理分配奖励
- 样本效率提升,收敛速度更快,优于均匀赋信方法
- 适合研究扩散模型与强化学习结合的学者
Flow-GRPO成功将强化学习应用于流模型,但对所有生成步骤采用统一奖励分配,忽略了扩散生成过程的时间结构:早期步骤决定内容与构图(低频结构),晚期步骤处理细节与纹理(高频细节)。仅根据最终图像分配统一奖励,可能无意中奖励了中间阶段表现不佳的生成步骤,尤其当错误在轨迹后期被修正时。本文提出Stepwise-Flow-GRPO,基于每步的奖励改进进行分步赋信。通过利用Tweedie公式获取中间奖励估计,并引入基于收益的优势函数,显著提升了样本效率和收敛速度。此外,我们设计了一种受DDIM启发的SDE,在保持随机性以支持策略梯度的同时,提升了奖励质量。
原文摘要 · Abstract (English)
Flow-GRPO successfully applies reinforcement learning to flow models, but uses uniform credit assignment across all steps. This ignores the temporal structure of diffusion generation: early steps determine composition and content (low-frequency structure), while late steps resolve details and textures (high-frequency details). Moreover, assigning uniform credit based solely on the final image can inadvertently reward suboptimal intermediate steps, especially when errors are corrected later in the diffusion trajectory. We propose Stepwise-Flow-GRPO, which assigns credit based on each step's reward improvement. By leveraging Tweedie's formula to obtain intermediate reward estimates and introducing gain-based advantages, our method achieves superior sample efficiency and faster convergence. We also introduce a DDIM-inspired SDE that improves reward quality while preserving stochasticity for policy gradients.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。