让生成过程的每一步都重要:时间感知的强化学习优化方法
TempFlow-GRPO: When Timing Matters for GRPO in Flow Models
- 在生成关键步骤设置分支点,精准分配奖励
- 根据各阶段探索潜力动态调整优化强度,早期重点学习
- 控制初始条件影响,分离探索贡献,适合追求高质量生成的开发者
近期基于流匹配的文生图模型已取得卓越质量,但其与强化学习结合进行人类偏好对齐的效果仍不理想,难以实现细粒度奖励优化。我们发现,现有方法中时间均匀性假设是主要障碍:稀疏终端奖励配合均匀信用分配无法捕捉生成过程中各步骤决策的重要性差异,导致探索效率低、收敛不佳。为此,我们提出 extbf{TempFlow-GRPO}(时间感知流式GRPO),一个能利用流生成内在时间结构的原理性框架。该方法引入三项创新:(i) 轨迹分支机制,在指定分支点集中随机性以提供过程奖励,实现精确信用分配,无需专门的中间奖励模型;(ii) 噪声感知加权策略,根据每个时间步的内在探索潜力调节策略优化,优先学习高影响早期阶段,保障后期稳定精修;(iii) 种子分组策略,控制初始化效应以隔离探索贡献。这些设计使模型具备时间感知优化能力,尊重生成动力学本质,在人类偏好对齐与文生图基准上达到当前最优性能。
原文摘要 · Abstract (English)
Recent flow matching models for text-to-image generation have achieved remarkable quality, yet their integration with reinforcement learning for human preference alignment remains suboptimal, hindering fine-grained reward-based optimization. We observe that the key impediment to effective GRPO training of flow models is the temporal uniformity assumption in existing approaches: sparse terminal rewards with uniform credit assignment fail to capture the varying criticality of decisions across generation timesteps, resulting in inefficient exploration and suboptimal convergence. To remedy this shortcoming, we introduce \textbf{TempFlow-GRPO} (Temporal Flow GRPO), a principled GRPO framework that captures and exploits the temporal structure inherent in flow-based generation. TempFlow-GRPO introduces three key innovations: (i) a trajectory branching mechanism that provides process rewards by concentrating stochasticity at designated branching points, enabling precise credit assignment without requiring specialized intermediate reward models; (ii) a noise-aware weighting scheme that modulates policy optimization according to the intrinsic exploration potential of each timestep, prioritizing learning during high-impact early stages while ensuring stable refinement in later phases; and (iii) a seed group strategy that controls for initialization effects to isolate exploration contributions. These innovations endow the model with temporally-aware optimization that respects the underlying generative dynamics, leading to state-of-the-art performance in human preference alignment and text-to-image benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。