用局部奖励提升视频生成质量,解决细节错误问题。
Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models
- 引入像素块级奖励模型捕捉局部缺陷
- 比基线方法在两项评估中显著提升质量
- 适合追求高精度视频生成的研究者
扩散模型(DMs)显著提升了文本到视频生成模型(VGMs)的质量。然而,当前VGM优化主要关注视频整体质量,忽视局部错误,导致生成效果不理想。为此,我们提出一种后训练策略HALO,显式引入由图像块奖励模型提供的局部反馈,为高级VGM优化提供详尽的训练信号。为构建高效块奖励模型,我们通过蒸馏GPT-4o持续训练视频奖励模型,提升训练效率并确保视频与块级奖励分布的一致性。此外,为协同整合块奖励,我们设计了粒度化的直接偏好优化(Gran-DPO)算法,实现块与视频奖励在优化过程中的协同使用。实验表明,我们的块奖励模型与人工标注高度一致,HALO在两种评估方式下均显著优于基线。进一步定量分析证实了块级缺陷的存在,且本方法能有效缓解该问题。
原文摘要 · Abstract (English)
The emergence of diffusion models (DMs) has significantly improved the quality of text-to-video generation models (VGMs). However, current VGM optimization primarily emphasizes the global quality of videos, overlooking localized errors, which leads to suboptimal generation capabilities. To address this issue, we propose a post-training strategy for VGMs, HALO, which explicitly incorporates local feedback from a patch reward model, providing detailed and comprehensive training signals with the video reward model for advanced VGM optimization. To develop an effective patch reward model, we distill GPT-4o to continuously train our video reward model, which enhances training efficiency and ensures consistency between video and patch reward distributions. Furthermore, to harmoniously integrate patch rewards into VGM optimization, we introduce a granular DPO (Gran-DPO) algorithm for DMs, allowing collaborative use of both patch and video rewards during the optimization process. Experimental results indicate that our patch reward model aligns well with human annotations and HALO substantially outperforms the baselines across two evaluation methods. Further experiments quantitatively prove the existence of patch defects, and our proposed method could effectively alleviate this issue.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。