通过数据与奖励信号优化扩散模型,提升文本生成视频质量。
T2V-Turbo-v2: Enhancing Video Generation Model Post-Training through Data, Reward, and Conditional Guidance Design
- 用高质量数据和奖励反馈指导一致性蒸馏过程。
- 在VBench上达到85.13总分,超越商用模型如Gen-3。
- 设计运动引导机制,显著改善生成视频的动作流畅性。
本文聚焦于基于扩散的文本到视频(T2V)模型在后训练阶段的增强,通过从预训练的T2V模型中提炼出一个高性能的一致性模型。提出的T2V-Turbo-v2方法将高质量训练数据、奖励模型反馈和条件引导等多路监督信号融入一致性蒸馏流程。通过全面的消融实验,我们强调了针对特定学习目标定制数据集的重要性,以及从多样奖励模型中学习对提升视觉质量和文本-视频对齐的有效性。此外,我们探索了条件引导策略的广阔设计空间,核心在于设计有效的能量函数以增强教师ODE求解器。通过从训练数据中提取运动引导并融入ODE求解器,显著提升了生成视频的运动质量,在VBench和T2V-CompBench上的运动相关指标均有提升。实证表明,T2V-Turbo-v2在VBench上取得85.13的总分,创下新纪录,超越了Gen-3和Kling等专有系统。
原文摘要 · Abstract (English)
In this paper, we focus on enhancing a diffusion-based text-to-video (T2V) model during the post-training phase by distilling a highly capable consistency model from a pretrained T2V model. Our proposed method, T2V-Turbo-v2, introduces a significant advancement by integrating various supervision signals, including high-quality training data, reward model feedback, and conditional guidance, into the consistency distillation process. Through comprehensive ablation studies, we highlight the crucial importance of tailoring datasets to specific learning objectives and the effectiveness of learning from diverse reward models for enhancing both the visual quality and text-video alignment. Additionally, we highlight the vast design space of conditional guidance strategies, which centers on designing an effective energy function to augment the teacher ODE solver. We demonstrate the potential of this approach by extracting motion guidance from the training datasets and incorporating it into the ODE solver, showcasing its effectiveness in improving the motion quality of the generated videos with the improved motion-related metrics from VBench and T2V-CompBench. Empirically, our T2V-Turbo-v2 establishes a new state-of-the-art result on VBench, with a Total score of 85.13, surpassing proprietary systems such as Gen-3 and Kling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。