arXiv:2508.06082cs.CV2025-08AAAI被引 6

让视频生成只需几步就能保持高质量,靠的是轨迹与分布双重对齐。

SwiftVideo: A Unified Framework for Few-Step Video Generation through Trajectory-Distribution Alignment

  • 通过连续时间一致性蒸馏,精准保留微分方程轨迹。
  • 在少步生成下,比现有方法显著提升视频质量,尤其在OpenVid-1M上表现优异。
  • 适合需要快速生成高质量视频的场景,如实时应用或资源受限环境。

基于扩散或流模型的视频生成已取得显著进展,但需多步迭代采样,计算开销大。许多仅依赖轨迹保持或分布匹配的蒸馏方法虽加速生成,但在少步设置下易性能崩溃或引入伪影。为此,我们提出 extbf{ extit{SwiftVideo}},一个统一且稳定的蒸馏框架,融合轨迹保持与分布匹配优势。通过引入连续时间一致性蒸馏,确保微分方程轨迹精确保留;进一步提出双视角对齐:合成数据与真实数据间的分布对齐,以及不同推理步长间的轨迹对齐。该方法在大幅减少推理步数的同时,仍保持高质量视频生成。在OpenVid-1M基准上的定量评估表明,本方法在少步视频生成中显著优于现有方法。

原文摘要 · Abstract (English)

Diffusion-based or flow-based models have achieved significant progress in video synthesis but require multiple iterative sampling steps, which incurs substantial computational overhead. While many distillation methods that are solely based on trajectory-preserving or distribution-matching have been developed to accelerate video generation models, these approaches often suffer from performance breakdown or increased artifacts under few-step settings. To address these limitations, we propose \textbf{\emph{SwiftVideo}}, a unified and stable distillation framework that combines the advantages of trajectory-preserving and distribution-matching strategies. Our approach introduces continuous-time consistency distillation to ensure precise preservation of ODE trajectories. Subsequently, we propose a dual-perspective alignment that includes distribution alignment between synthetic and real data along with trajectory alignment across different inference steps. Our method maintains high-quality video generation while substantially reducing the number of inference steps. Quantitative evaluations on the OpenVid-1M benchmark demonstrate that our method significantly outperforms existing approaches in few-step video generation.

视频生成扩散模型少步生成蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。