arXiv:2512.12387cs.LG2025-12被引 4

改进图像生成的奖励机制,让模型更懂何时该重视哪一步。

Anchoring Values in Temporal and Group Dimensions for Flow Matching Model Alignment

  • 用动态价值估计替代单一终点奖励,精准分配每一步生成的贡献。
  • 在三个基准上实现顶尖图像质量,且任务准确率显著提升。
  • 适合研究扩散模型对齐与强化学习融合的学者参考。

组相对策略优化(GRPO)在提升大语言模型对齐能力方面表现优异,但现有针对基于流匹配的图像生成方法的适配存在根本性矛盾:其核心原则与视觉合成过程动态不匹配。这导致两大局限:(i) 在所有时间步均匀使用稀疏终端奖励,破坏了时间信用分配,忽视了从早期结构形成到后期调优各阶段的重要性差异;(ii) 仅依赖组内相对奖励使优化信号随训练收敛而衰减,导致奖励多样性耗尽后优化停滞。为此,我们提出值锚定组策略优化(VGPO),重构时空维度上的价值估计。具体地,将稀疏终端奖励转化为密集、过程感知的价值估计,通过建模每个生成阶段的期望累积奖励实现精确信用分配。此外,用引入绝对值的新式过程增强组归一化,即使奖励多样性下降也能维持稳定优化信号。在三个基准上的实验表明,VGPO在实现顶尖图像质量的同时,显著提升任务特定准确率,有效缓解奖励欺骗问题。

原文摘要 · Abstract (English)

Group Relative Policy Optimization (GRPO) has proven highly effective in enhancing the alignment capabilities of Large Language Models (LLMs). However, current adaptations of GRPO for the flow matching-based image generation neglect a foundational conflict between its core principles and the distinct dynamics of the visual synthesis process. This mismatch leads to two key limitations: (i) Uniformly applying a sparse terminal reward across all timesteps impairs temporal credit assignment, ignoring the differing criticality of generation phases from early structure formation to late-stage tuning. (ii) Exclusive reliance on relative, intra-group rewards causes the optimization signal to fade as training converges, leading to the optimization stagnation when reward diversity is entirely depleted. To address these limitations, we propose Value-Anchored Group Policy Optimization (VGPO), a framework that redefines value estimation across both temporal and group dimensions. Specifically, VGPO transforms the sparse terminal reward into dense, process-aware value estimates, enabling precise credit assignment by modeling the expected cumulative reward at each generative stage. Furthermore, VGPO replaces standard group normalization with a novel process enhanced by absolute values to maintain a stable optimization signal even as reward diversity declines. Extensive experiments on three benchmarks demonstrate that VGPO achieves state-of-the-art image quality while simultaneously improving task-specific accuracy, effectively mitigating reward hacking. Project webpage: https://yawen-shao.github.io/VGPO/.

图像生成强化学习流匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。