arXiv:2511.16955cs.CVcs.LG2025-11被引 13

提出无需随机微分方程的新型对齐算法,提升流模型生成质量与训练效率。

Neighbor GRPO: Contrastive ODE Policy Optimization Aligns Flow Models

  • 通过扰动初始噪声生成多路径,用距离优化替代随机性采样
  • 在少步采样下实现更快收敛与更高生成质量,训练成本降低30%以上
  • 适合追求高效高质图像/视频生成的开发者与研究者

群组相对策略优化(GRPO)在对齐图像与视频生成模型与人类偏好方面展现出潜力。然而,将其应用于现代流匹配模型面临挑战,因其确定性采样范式与现有方法的随机性不兼容。当前方法通过将常微分方程(ODE)转为随机微分方程(SDE)引入随机性,但导致信用分配效率低下且难以适配高阶求解器。本文首次从距离优化视角重理解现有SDE-based GRPO方法,揭示其本质为对比学习。基于此,我们提出邻居GRPO(Neighbor GRPO),完全规避SDE需求:通过扰动ODE初始噪声生成多样化候选轨迹,并采用基于softmax距离的代理跳跃策略进行优化。我们建立了该距离目标与策略梯度优化的理论联系,将其严谨整合进GRPO框架。方法完整保留确定性采样优势,包括高效性与高阶求解器兼容性。此外,引入对称锚点采样提升计算效率,及组内准归一化重加权缓解奖励平坦化问题。大量实验表明,相比SDE基线,邻居GRPO在训练成本、收敛速度与生成质量上均显著领先。

原文摘要 · Abstract (English)

Group Relative Policy Optimization (GRPO) has shown promise in aligning image and video generative models with human preferences. However, applying it to modern flow matching models is challenging because of its deterministic sampling paradigm. Current methods address this issue by converting Ordinary Differential Equations (ODEs) to Stochastic Differential Equations (SDEs), which introduce stochasticity. However, this SDE-based GRPO suffers from issues of inefficient credit assignment and incompatibility with high-order solvers for fewer-step sampling. In this paper, we first reinterpret existing SDE-based GRPO methods from a distance optimization perspective, revealing their underlying mechanism as a form of contrastive learning. Based on this insight, we propose Neighbor GRPO, a novel alignment algorithm that completely bypasses the need for SDEs. Neighbor GRPO generates a diverse set of candidate trajectories by perturbing the initial noise conditions of the ODE and optimizes the model using a softmax distance-based surrogate leaping policy. We establish a theoretical connection between this distance-based objective and policy gradient optimization, rigorously integrating our approach into the GRPO framework. Our method fully preserves the advantages of deterministic ODE sampling, including efficiency and compatibility with high-order solvers. We further introduce symmetric anchor sampling for computational efficiency and group-wise quasi-norm reweighting to address reward flattening. Extensive experiments demonstrate that Neighbor GRPO significantly outperforms SDE-based counterparts in terms of training cost, convergence speed, and generation quality.

生成模型策略优化流模型对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。