arXiv:2503.19462cs.CV2025-03被引 12

用合成数据加速视频生成,8.5倍提速且画质更高

AccVideo: Accelerating Video Diffusion Model with Synthetic Dataset

  • 用预训练模型生成合成去噪轨迹数据
  • 仅需8步即可生成5秒720x1280高清视频
  • 适合需要高速高质量视频生成的场景

扩散模型在视频生成领域取得显著进展,但其迭代去噪机制导致生成速度慢、计算成本高。本文分析现有扩散蒸馏方法的挑战,提出AccVideo方法,通过预训练视频扩散模型生成多条有效去噪轨迹作为合成数据集,避免蒸馏过程中的无效数据点。基于该合成数据集,设计基于轨迹的少步引导策略,利用关键去噪点学习噪声到视频的映射,实现少步生成。同时,因合成数据捕捉各扩散步的数据分布,引入对抗训练策略对齐学生模型输出分布,提升视频质量。大量实验表明,本方法相较教师模型生成速度提升8.5倍,且可生成5秒、720x1280、24fps的高分辨率视频,优于现有加速方法。

原文摘要 · Abstract (English)

Diffusion models have achieved remarkable progress in the field of video generation. However, their iterative denoising nature requires a large number of inference steps to generate a video, which is slow and computationally expensive. In this paper, we begin with a detailed analysis of the challenges present in existing diffusion distillation methods and propose a novel efficient method, namely AccVideo, to reduce the inference steps for accelerating video diffusion models with synthetic dataset. We leverage the pretrained video diffusion model to generate multiple valid denoising trajectories as our synthetic dataset, which eliminates the use of useless data points during distillation. Based on the synthetic dataset, we design a trajectory-based few-step guidance that utilizes key data points from the denoising trajectories to learn the noise-to-video mapping, enabling video generation in fewer steps. Furthermore, since the synthetic dataset captures the data distribution at each diffusion timestep, we introduce an adversarial training strategy to align the output distribution of the student model with that of our synthetic dataset, thereby enhancing the video quality. Extensive experiments demonstrate that our model achieves 8.5x improvements in generation speed compared to the teacher model while maintaining comparable performance. Compared to previous accelerating methods, our approach is capable of generating videos with higher quality and resolution, i.e., 5-seconds, 720x1280, 24fps.

视频生成扩散模型加速合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。