视频生成中优化奖励信号,防止模型钻空子、过拟合。
Rethinking Reward Signals in Video GRPO: When Scores Become Targets
- 用组件评估+组内稀疏性重构多维度奖励信号
- 自适应降低饱和组件权重,保持优化方向有效
- 提升视频质量、运动连贯性和图文对齐,适合视频生成研究者
Group Relative Policy Optimization (GRPO) 通过群体比较实现稳定且偏好导向的后训练视频生成。然而,GRPO 直接优化由奖励诱导的优势,在持续优化下,奖励分数可能偏离真实视频质量,违背 Goodhart's Law。这导致两个常见问题:(i) 复合目标下的捷径优化,(ii) 提示组内的奖励饱和。为此,我们提出 TaRoS——一种针对视频生成 GRPO 的目标鲁棒奖励信号框架。TaRoS 结合组件级性能评估与组内稀疏性,组织多方面奖励以对齐优化目标。同时,它自适应降低出现饱和的组件权重,从而保留有效优化方向并减少冗余。该方法维持了有意义的优化路径和组内排序分离,防止奖励劫持,带来更可靠的策略更新。大量实验表明,相比强基线,其在视觉保真度、运动连贯性和文本-视频对齐上均有持续提升。
原文摘要 · Abstract (English)
Group Relative Policy Optimization (GRPO) enables stable and preference-oriented updates via group-wise comparisons for post-training video generation. However, GRPO directly optimizes reward-induced advantages. Under sustained optimization, the reward score can lose fidelity as a proxy for true video quality, consistent with the phenomenon described by Goodhart's Law. This leads to two recurring issues: (i) shortcut-driven optimization under composite objectives and (ii) reward saturation within prompt groups. To address these issues, we introduce TaRoS, a Target-Robust Reward Signaling framework for Video generation GRPO. TaRoS leverages component level performance assessment together with intra-group sparsity to organize multi-aspect rewards towards optimization objectives. In addition, it adaptively downweights components that exhibit saturation, thereby preserving effective optimization directions and mitigating redundancy. This maintains meaningful optimization directions and preserves within-group ranking separation, thereby preventing reward hacking and leading to more reliable policy updates. Extensive experiments show consistent improvements in visual fidelity, motion coherence, and text-video alignment over strong baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。