系统梳理视频生成从生成对抗网络到多模态大模型的演进历程。
Evolution of Video Generative Foundations
- 按技术发展脉络,梳理视频生成从GAN到扩散模型再到自回归与多模态融合的演进路径。
- 对比分析各类模型在时序连贯性与语义丰富性上的优势与局限。
- 适合关注视频生成前沿、世界模型构建及多模态应用的研究者阅读。
人工智能生成内容(AIGC)的迅猛发展推动了视频生成技术的革新,催生了包括OpenAI Sora、Google Veo3、字节跳动Seedance在内的商业先行者,以及Wan和HunyuanVideo等开源模型,它们能生成时空连贯且语义丰富的视频。这些进展为构建模拟真实世界动态的“世界模型”铺平道路,应用场景涵盖娱乐、教育与虚拟现实。然而,现有综述多聚焦于特定技术(如生成对抗网络GAN、扩散模型)或具体任务(如视频编辑),缺乏对视频生成领域整体演进的全面审视,尤其忽视了自回归(AR)模型与多模态信息融合的发展。为此,本文首次系统回顾视频生成技术的演进:从早期的GAN,到主导的扩散模型,再到新兴的AR模型与多模态技术。深入分析其基础原理、关键进展及优劣对比;探讨多模态视频生成趋势,强调融合多种数据类型以增强上下文感知能力;并结合历史发展与当前创新,为未来研究提供指导,涵盖虚拟/增强现实、个性化教育、自动驾驶仿真、数字娱乐及高级世界模型等方向。更多信息详见项目页:https://github.com/sjtuplayer/Awesome-Video-Foundations。
原文摘要 · Abstract (English)
The rapid advancement of Artificial Intelligence Generated Content (AIGC) has revolutionized video generation, enabling systems ranging from proprietary pioneers like OpenAI's Sora, Google's Veo3, and Bytedance's Seedance to powerful open-source contenders like Wan and HunyuanVideo to synthesize temporally coherent and semantically rich videos. These advancements pave the way for building "world models" that simulate real-world dynamics, with applications spanning entertainment, education, and virtual reality. However, existing reviews on video generation often focus on narrow technical fields, e.g., Generative Adversarial Networks (GAN) and diffusion models, or specific tasks (e. g., video editing), lacking a comprehensive perspective on the field's evolution, especially regarding Auto-Regressive (AR) models and integration of multimodal information. To address these gaps, this survey firstly provides a systematic review of the development of video generation technology, tracing its evolution from early GANs to dominant diffusion models, and further to emerging AR-based and multimodal techniques. We conduct an in-depth analysis of the foundational principles, key advancements, and comparative strengths/limitations. Then, we explore emerging trends in multimodal video generation, emphasizing the integration of diverse data types to enhance contextual awareness. Finally, by bridging historical developments and contemporary innovations, this survey offers insights to guide future research in video generation and its applications, including virtual/augmented reality, personalized education, autonomous driving simulations, digital entertainment, and advanced world models, in this rapidly evolving field. For more details, please refer to the project at https://github.com/sjtuplayer/Awesome-Video-Foundations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。