用人类反馈提升视频生成质量,解决动作不顺和图文不符问题。
Improving Video Generation with Human Feedback

- 构建大规模人类偏好数据集,设计多维度视频评分模型VideoReward。
- 提出Flow-DPO等三种对齐算法,显著改善生成视频流畅度与文本一致性。
- 支持用户在推理时自定义权重,满足个性化视频生成需求。
视频生成虽借助修正流技术取得显著进展,但仍存在运动不连贯、视频与提示不一致等问题。本文构建了一个面向现代视频生成模型的大规模人类偏好数据集,包含多维度成对标注。提出VideoReward多维视频奖励模型,并研究标注方式与设计选择对其性能的影响。基于统一强化学习框架(最大化奖励并加入KL正则化),提出三种适用于流模型的对齐算法:两种训练阶段方法——直接偏好优化(Flow-DPO)与奖励加权回归(Flow-RWR),以及一种推理阶段技术——流噪声引导(Flow-NRG),可直接对噪声视频施加奖励引导。实验表明,VideoReward显著优于现有奖励模型,Flow-DPO在性能上超越Flow-RWR与监督微调方法。此外,Flow-NRG支持用户在推理时为多个目标分配自定义权重,满足个性化视频质量需求。
原文摘要 · Abstract (English)
Video generation has achieved significant advances through rectified flow techniques, but issues like unsmooth motion and misalignment between videos and prompts persist. In this work, we develop a systematic pipeline that harnesses human feedback to mitigate these problems and refine the video generation model. Specifically, we begin by constructing a large-scale human preference dataset focused on modern video generation models, incorporating pairwise annotations across multi-dimensions. We then introduce VideoReward, a multi-dimensional video reward model, and examine how annotations and various design choices impact its rewarding efficacy. From a unified reinforcement learning perspective aimed at maximizing reward with KL regularization, we introduce three alignment algorithms for flow-based models. These include two training-time strategies: direct preference optimization for flow (Flow-DPO) and reward weighted regression for flow (Flow-RWR), and an inference-time technique, Flow-NRG, which applies reward guidance directly to noisy videos. Experimental results indicate that VideoReward significantly outperforms existing reward models, and Flow-DPO demonstrates superior performance compared to both Flow-RWR and supervised fine-tuning methods. Additionally, Flow-NRG lets users assign custom weights to multiple objectives during inference, meeting personalized video quality needs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。