用人类偏好优化视频生成,提升真实感与美感。
HuViDPO:Enhancing Video Generation through Direct Preference Optimization for Human-Centric Alignment
- 首次将直接偏好优化引入文本到视频生成,构建精准损失函数。
- 自建小规模动作类别偏好数据集,降低训练成本并提升画质。
- 基于首帧引导+稀疏因果注意力,增强生成灵活性与质量。
随着AIGC技术的快速发展,基于扩散模型的文本到图像(T2I)和文本到视频(T2V)生成取得了显著进展。近年来,已有研究将直接偏好优化(DPO)应用于T2I任务,显著提升了生成图像的人类偏好契合度。然而,现有T2V生成方法缺乏明确的损失函数指导,难以通过DPO策略对齐人类偏好;同时,成对视频偏好数据稀缺,制约了模型训练。此外,训练数据不足可能导致生成视频灵活性差、质量不高。针对这些问题,本文提出三项改进:1)首次将DPO策略引入T2V任务,推导出结构化损失函数,利用人类反馈对齐视频生成与人类偏好,提出HuViDPO方法;2)为每类动作构建小规模人类偏好数据集,并用于微调,有效提升生成视频的美学质量,同时降低训练开销;3)采用首帧条件策略,利用首帧丰富信息引导后续帧生成,提升生成灵活性;同时引入稀疏因果注意力机制,进一步提升生成质量。更多细节与示例见官网:https://tankowa.github.io/HuViDPO.github.io/
原文摘要 · Abstract (English)
With the rapid development of AIGC technology, significant progress has been made in diffusion model-based technologies for text-to-image (T2I) and text-to-video (T2V). In recent years, a few studies have introduced the strategy of Direct Preference Optimization (DPO) into T2I tasks, significantly enhancing human preferences in generated images. However, existing T2V generation methods lack a well-formed pipeline with exact loss function to guide the alignment of generated videos with human preferences using DPO strategies. Additionally, challenges such as the scarcity of paired video preference data hinder effective model training. At the same time, the lack of training datasets poses a risk of insufficient flexibility and poor video generation quality in the generated videos. Based on those problems, our work proposes three targeted solutions in sequence. 1) Our work is the first to introduce the DPO strategy into the T2V tasks. By deriving a carefully structured loss function, we utilize human feedback to align video generation with human preferences. We refer to this new method as HuViDPO. 2) Our work constructs small-scale human preference datasets for each action category and fine-tune this model, improving the aesthetic quality of the generated videos while reducing training costs. 3) We adopt a First-Frame-Conditioned strategy, leveraging the rich in formation from the first frame to guide the generation of subsequent frames, enhancing flexibility in video generation. At the same time, we employ a SparseCausal Attention mechanism to enhance the quality of the generated videos.More details and examples can be accessed on our website: https://tankowa.github.io/HuViDPO. github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。