让视频生成更符合用户偏好,兼顾画质与语义匹配。
VideoDPO: Omni-Preference Alignment for Video Diffusion Generation

- 用双维度评分机制综合评估视频画质与文本语义匹配度。
- 自动构建偏好数据对并按评分重加权,显著提升对齐效果。
- 适合关注视频生成质量与用户意图一致性的研究者。
生成式扩散模型的进展极大推动了文本到视频的生成能力。尽管在大规模多样数据上训练的文本到视频模型能生成丰富结果,但这些输出常偏离用户偏好,凸显了对预训练模型进行偏好对齐的必要性。虽然直接偏好优化(DPO)已在语言和图像生成中取得显著成效,本文首次将其适配至视频扩散模型,并提出VideoDPO流程,通过多项关键调整实现突破。不同于以往仅关注视觉质量或文本-视频语义对齐的图像对齐方法,我们全面考虑两个维度,构建综合评分体系,称为OmniScore。设计了一套基于OmniScore自动生成偏好数据对的流程,并发现按评分重加权这些数据对可显著影响整体偏好对齐效果。实验表明,该方法在视觉质量和语义对齐方面均实现显著提升,确保无任何偏好维度被忽视。代码与数据将公开于https://videodpo.github.io/。
原文摘要 · Abstract (English)
Recent progress in generative diffusion models has greatly advanced text-to-video generation. While text-to-video models trained on large-scale, diverse datasets can produce varied outputs, these generations often deviate from user preferences, highlighting the need for preference alignment on pre-trained models. Although Direct Preference Optimization (DPO) has demonstrated significant improvements in language and image generation, we pioneer its adaptation to video diffusion models and propose a VideoDPO pipeline by making several key adjustments. Unlike previous image alignment methods that focus solely on either (i) visual quality or (ii) semantic alignment between text and videos, we comprehensively consider both dimensions and construct a preference score accordingly, which we term the OmniScore. We design a pipeline to automatically collect preference pair data based on the proposed OmniScore and discover that re-weighting these pairs based on the score significantly impacts overall preference alignment. Our experiments demonstrate substantial improvements in both visual quality and semantic alignment, ensuring that no preference aspect is neglected. Code and data will be shared at https://videodpo.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。