用视频质量评估模型优化扩散模型生成视频,解决画质差和闪烁问题。
OnlineVPO: Align Video Diffusion Model with Online Video-Centric Preference Optimization
- 用视频质量评估模型替代图像奖励模型,提升反馈与人类感知对齐度。
- 提出在线偏好优化算法,支持高分辨率长视频高效训练。
- 适合追求高质量视频生成的开发者和研究者使用。
视频扩散模型(VDM)在文本到视频生成任务中表现卓越,但仍存在画质下降和闪烁伪影问题。现有方法虽引入偏好学习利用人类反馈改进生成质量,但多沿用图像领域范式,未深入探索视频特性。本文从反馈来源与优化方法两方面重新审视视频偏好学习,提出面向VDM的OnlineVPO框架。实验发现,图像级奖励模型因模态差异难以提供与人类一致的视频偏好信号;而视频质量评估(VQA)模型更贴近人类感知,可作为代理反馈。基于此,采用VQA模型提供更对齐的视频反馈。在优化方法上,设计适用于VDM的在线深度偏好优化(DPO)算法,兼具高可扩展性,并通过在线偏好生成与课程式偏好更新缓解离策略学习导致的优化不足。在开源VDM上的大量实验表明,OnlineVPO是一种简单、有效且可扩展的视频扩散模型偏好学习方法。
原文摘要 · Abstract (English)
Video diffusion models (VDMs) have demonstrated remarkable capabilities in text-to-video (T2V) generation. Despite their success, VDMs still suffer from degraded image quality and flickering artifacts. To address these issues, some approaches have introduced preference learning to exploit human feedback to enhance the video generation. However, these methods primarily adopt the routine in the image domain without an in-depth investigation into video-specific preference optimization. In this paper, we reexamine the design of the video preference learning from two key aspects: feedback source and feedback tuning methodology, and present OnlineVPO, a more efficient preference learning framework tailored specifically for VDMs. On the feedback source, we found that the image-level reward model commonly used in existing methods fails to provide a human-aligned video preference signal due to the modality gap. In contrast, video quality assessment (VQA) models show superior alignment with human perception of video quality. Building on this insight, we propose leveraging VQA models as a proxy of humans to provide more modality-aligned feedback for VDMs. Regarding the preference tuning methodology, we introduce an online DPO algorithm tailored for VDMs. It not only enjoys the benefits of superior scalability in optimizing videos with higher resolution and longer duration compared with the existing method, but also mitigates the insufficient optimization issue caused by off-policy learning via online preference generation and curriculum preference update designs. Extensive experiments on the open-source video-diffusion model demonstrate OnlineVPO as a simple yet effective and, more importantly, scalable preference learning algorithm for video diffusion models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。