系统梳理视频生成模型的高效化方法,助力落地应用
Efficient Video Diffusion Models: Advancements and Challenges

- 按四类范式分类:步数压缩、高效注意力、模型压缩、缓存优化
- 聚焦减少计算量与每步开销,提升推理效率
- 适合关注视频生成部署与性能优化的研究者
视频扩散模型已成为高保真生成视频的主流范式,但实际部署受限于高昂的推理成本。相较于图像生成,视频合成因时空标记增长与迭代去噪带来更大计算与内存压力,导致注意力与内存访问成为现实场景中的主要瓶颈。本文系统性地回顾了高效视频扩散模型,提出统一分类框架,将现有方法归纳为四类核心范式:步数蒸馏、高效注意力、模型压缩与缓存/轨迹优化。基于该分类,分析各类方法的算法趋势,探讨不同设计如何针对两大核心目标——减少函数评估次数与降低每步开销。最后,讨论开放挑战与未来方向,包括复合加速下的质量保持、软硬件协同设计、鲁棒的实时长时生成以及标准化评估基础设施。据我们所知,这是首个关于高效视频扩散模型的全面综述,为研究者与工程师提供领域结构化概览与新兴方向指引。
原文摘要 · Abstract (English)
Video diffusion models have rapidly become the dominant paradigm for high-fidelity generative video synthesis, but their practical deployment remains constrained by severe inference costs. Compared with image generation, video synthesis compounds computation across spatial-temporal token growth and iterative denoising, making attention and memory traffic major bottlenecks in real-world settings. This survey provides a systematic and deployment-oriented review of efficient video diffusion models. We propose a unified categorization that organizes existing methods into four classes of main paradigms, including step distillation, efficient attention, model compression, and cache/trajectory optimization. Building on this categorization, we respectively analyze algorithmic trends of these four paradigms and examine how different design choices target two core objectives: reducing the number of function evaluations and minimizing per-step overhead. Finally, we discuss open challenges and future directions, including quality preservation under composite acceleration, hardware-software co-design, robust real-time long-horizon generation, and open infrastructure for standardized evaluation. To the best of our knowledge, our work is the first comprehensive survey on efficient video diffusion models, offering researchers and engineers a structured overview of the field and its emerging research directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。