通过轨迹对齐提升图像到视频生成的奖励表现
TAGRPO: Boosting GRPO on Image-to-Video Generation with Direct Trajectory Alignment
- 在中间潜在空间直接对齐高奖励与低奖励轨迹
- 相比DanceGRPO,视频生成质量显著提升
- 适合需要高质量视频生成的研究者和开发者
近期研究证明将组相对策略优化(GRPO)整合到流匹配模型中,在文本到图像和文本到视频生成中效果显著。然而,我们发现直接将这些技术应用于图像到视频(I2V)模型时,难以持续提升奖励。为此,我们提出TAGRPO,一种受对比学习启发的鲁棒后训练框架。该方法基于同一初始噪声生成的轨迹提供更优优化指引的观察,提出一种新型GRPO损失,作用于中间潜在表示,鼓励与高奖励轨迹对齐,同时最大化与低奖励轨迹的距离。此外,引入回放缓存机制以增强多样性并降低计算开销。尽管结构简单,TAGRPO在I2V生成上显著优于DanceGRPO。
原文摘要 · Abstract (English)
Recent studies have demonstrated the efficacy of integrating Group Relative Policy Optimization (GRPO) into flow matching models, particularly for text-to-image and text-to-video generation. However, we find that directly applying these techniques to image-to-video (I2V) models often fails to yield consistent reward improvements. To address this limitation, we present TAGRPO, a robust post-training framework for I2V models inspired by contrastive learning. Our approach is grounded in the observation that rollout videos generated from identical initial noise provide superior guidance for optimization. Leveraging this insight, we propose a novel GRPO loss applied to intermediate latents, encouraging direct alignment with high-reward trajectories while maximizing distance from low-reward counterparts. Furthermore, we introduce a memory bank for rollout videos to enhance diversity and reduce computational overhead. Despite its simplicity, TAGRPO achieves significant improvements over DanceGRPO in I2V generation. The deliverables will be updated at https://tagrpo.github.io/ .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。