arXiv:2601.00423cs.LGcs.AI2026-01被引 14

通过提升采样步的熵值,让流模型强化学习更高效。

E-GRPO: High Entropy Steps Drive Effective Reinforcement Learning for Flow Models

  • 用熵感知策略优化采样步,合并低熵步形成高熵步。
  • 在多种奖励设置下,生成质量显著优于现有方法。
  • 适合做对齐人类偏好的文本生成或图像生成任务。

近期强化学习已提升流匹配模型在人类偏好对齐上的表现。尽管随机采样能探索去噪方向,但现有方法在多个去噪步上优化时面临稀疏且模糊的奖励信号问题。我们观察到,高熵步能实现更高效、有效的探索,而低熵步则导致生成结果差异小。为此,提出E-GRPO:一种熵感知的组相对策略优化方法,用于提升SDE采样步的熵值。由于多步随机性导致奖励信号模糊,我们特别将连续的低熵步合并为一个高熵步进行SDE采样,其余步采用ODE采样。在此基础上,引入多步组归一化优势,计算共享同一合并SDE去噪步样本间的组相对优势。在不同奖励设置下的实验表明该方法有效。

原文摘要 · Abstract (English)

Recent reinforcement learning has enhanced the flow matching models on human preference alignment. While stochastic sampling enables the exploration of denoising directions, existing methods which optimize over multiple denoising steps suffer from sparse and ambiguous reward signals. We observe that the high entropy steps enable more efficient and effective exploration while the low entropy steps result in undistinguished roll-outs. To this end, we propose E-GRPO, an entropy aware Group Relative Policy Optimization to increase the entropy of SDE sampling steps. Since the integration of stochastic differential equations suffer from ambiguous reward signals due to stochasticity from multiple steps, we specifically merge consecutive low entropy steps to formulate one high entropy step for SDE sampling, while applying ODE sampling on other steps. Building upon this, we introduce multi-step group normalized advantage, which computes group-relative advantages within samples sharing the same consolidated SDE denoising step. Experimental results on different reward settings have demonstrated the effectiveness of our methods.

强化学习流模型采样优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。