arXiv:2504.16081cs.CVcs.CL2025-04中稿 · TMLR综述被引 20

系统梳理视频扩散模型的原理、方法与应用,助力高效生成高质量视频。

Survey of Video Diffusion Models: Foundations, Implementations, and Applications

  • 构建视频扩散模型的分类体系,涵盖架构设计与优化策略。
  • 总结多类任务应用效果,如去噪、超分辨率等,提升生成质量。
  • 聚焦评估指标与工业落地技术,适合研究者与开发者参考。

近年来,扩散模型在视频生成领域取得突破,相比传统基于生成对抗网络的方法,在时间一致性与视觉质量方面表现更优。尽管该领域前景广阔,仍面临运动一致性、计算效率及伦理挑战。本文系统综述了基于扩散模型的视频生成技术,涵盖其发展历程、技术基础与实际应用。提出一套完整的分类体系,分析架构创新与优化策略,并探讨在低层视觉任务(如去噪、超分辨率)中的应用。此外,还研究了视频生成与视频表征学习、问答、检索等领域的协同关系。相较已有综述(Lei et al., 2024a;b; Melnik et al., 2024; Cao et al., 2023; Xing et al., 2024c)仅关注特定方向(如人物视频合成或长视频生成),本工作提供更全面、更新颖且更细致的视角,特别增设评估指标、产业解决方案与训练工程技巧章节。本综述为扩散模型与视频生成交叉领域的研究人员和实践者提供理论与实现双重指导。相关文献列表详见 https://github.com/Eyeline-Research/Survey-Video-Diffusion。

原文摘要 · Abstract (English)

Recent advances in diffusion models have revolutionized video generation, offering superior temporal consistency and visual quality compared to traditional generative adversarial networks-based approaches. While this emerging field shows tremendous promise in applications, it faces significant challenges in motion consistency, computational efficiency, and ethical considerations. This survey provides a comprehensive review of diffusion-based video generation, examining its evolution, technical foundations, and practical applications. We present a systematic taxonomy of current methodologies, analyze architectural innovations and optimization strategies, and investigate applications across low-level vision tasks such as denoising and super-resolution. Additionally, we explore the synergies between diffusionbased video generation and related domains, including video representation learning, question answering, and retrieval. Compared to the existing surveys (Lei et al., 2024a;b; Melnik et al., 2024; Cao et al., 2023; Xing et al., 2024c) which focus on specific aspects of video generation, such as human video synthesis (Lei et al., 2024a) or long-form content generation (Lei et al., 2024b), our work provides a broader, more updated, and more fine-grained perspective on diffusion-based approaches with a special section for evaluation metrics, industry solutions, and training engineering techniques in video generation. This survey serves as a foundational resource for researchers and practitioners working at the intersection of diffusion models and video generation, providing insights into both the theoretical frameworks and practical implementations that drive this rapidly evolving field. A structured list of related works involved in this survey is also available on https://github.com/Eyeline-Research/Survey-Video-Diffusion.

视频生成扩散模型综述视觉质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。