通过知识蒸馏与潜在奖励优化,实现10秒视频的高效生成。
DOLLAR: Few-Step Video Generation via Distillation and Latent Reward Optimization
- 结合变分得分蒸馏与一致性蒸馏,提升少步采样效果。
- 4步生成模型在VBench上达82.57分,超越教师模型和多个基线。
- 支持非可微奖励优化,适合需定制化评估的生成任务。
扩散概率模型在视频生成中取得显著进展,但其计算效率受限于大量采样步骤。减少采样步数通常会损害视频质量或多样性。本文提出一种结合变分得分蒸馏与一致性蒸馏的知识蒸馏方法,实现少步视频生成,同时保持高质量与高多样性。此外,我们提出一种潜在奖励模型微调策略,可根据任意指定奖励指标进一步提升生成性能,该方法降低内存占用,且不要求奖励函数可微。该方法在10秒视频(128帧,12 FPS)的少步生成中达到当前最优表现。蒸馏后的学生模型在VBench上获得82.57分,优于教师模型及基线模型Gen-3、T2V-Turbo和Kling。单步蒸馏可使教师模型的扩散采样加速高达278.6倍,实现近实时生成。人类评估进一步验证了4步学生模型相比使用50步DDIM采样的教师模型具有更优表现。
原文摘要 · Abstract (English)
Diffusion probabilistic models have shown significant progress in video generation; however, their computational efficiency is limited by the large number of sampling steps required. Reducing sampling steps often compromises video quality or generation diversity. In this work, we introduce a distillation method that combines variational score distillation and consistency distillation to achieve few-step video generation, maintaining both high quality and diversity. We also propose a latent reward model fine-tuning approach to further enhance video generation performance according to any specified reward metric. This approach reduces memory usage and does not require the reward to be differentiable. Our method demonstrates state-of-the-art performance in few-step generation for 10-second videos (128 frames at 12 FPS). The distilled student model achieves a score of 82.57 on VBench, surpassing the teacher model as well as baseline models Gen-3, T2V-Turbo, and Kling. One-step distillation accelerates the teacher model's diffusion sampling by up to 278.6 times, enabling near real-time generation. Human evaluations further validate the superior performance of our 4-step student models compared to teacher model using 50-step DDIM sampling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。