arXiv:2510.01982cs.LGcs.CV2025-10被引 16

提出细粒度强化学习框架,提升流模型对人类偏好的精准对齐。

Fine-Grained GRPO for Precise Preference Alignment in Flow Models

  • 通过奇异随机采样实现分步探索,增强奖励分配精度。
  • 多粒度优势融合使采样轨迹评估更鲁棒,优于现有基线。
  • 适用于需高精度偏好对齐的生成任务,如文本到图像生成。

将在线强化学习引入扩散与流模型,已成为对齐模型行为与人类偏好的有力范式。通过在去噪阶段利用随机微分方程(SDE)进行随机采样,这些模型可探索多种去噪路径,提升强化学习的探索能力。然而,由于奖励信号稀疏且狭窄,现有方法在偏好对齐上仍表现不佳。为此,我们提出一种新框架——细粒度GRPO(G²RPO),实现流模型强化学习中采样方向的精细、全面评估。具体地,提出奇异随机采样机制,在保证注入噪声与奖励信号强相关的同时支持分步随机探索,提升每一步SDE扰动的信用分配精度。此外,为缓解固定粒度去噪带来的偏差,设计多粒度优势融合模块,聚合多尺度扩散过程中的优势值,实现更稳健的采样轨迹评估。在多种奖励模型(包括域内与域外设置)上的大量实验表明,本方法显著优于现有基于流的GRPO基线,验证了其有效性与泛化能力。

原文摘要 · Abstract (English)

The incorporation of online reinforcement learning (RL) into diffusion and flow-based generative models has recently gained attention as a powerful paradigm for aligning model behavior with human preferences. By leveraging stochastic sampling via Stochastic Differential Equations (SDEs) during the denoising phase, these models can explore a variety of denoising trajectories, enhancing the exploratory capacity of RL. However, despite their ability to discover potentially high-reward samples, current approaches often struggle to effectively align with preferences due to the sparsity and narrowness of reward feedback. To overcome this limitation, we introduce a novel framework called Granular-GRPO (G$^2$RPO), which enables fine-grained and comprehensive evaluation of sampling directions in the RL training of flow models. Specifically, we propose a Singular Stochastic Sampling mechanism that supports step-wise stochastic exploration while ensuring strong correlation between injected noise and reward signals, enabling more accurate credit assignment to each SDE perturbation. Additionally, to mitigate the bias introduced by fixed-granularity denoising, we design a Multi-Granularity Advantage Integration module that aggregates advantages computed across multiple diffusion scales, resulting in a more robust and holistic assessment of sampling trajectories. Extensive experiments on various reward models, including both in-domain and out-of-domain settings, demonstrate that our G$^2$RPO outperforms existing flow-based GRPO baselines, highlighting its effectiveness and generalization capability.

强化学习流模型偏好对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。