arXiv:2412.18688cs.CVcs.AI2024-12被引 10

综述长视频生成前沿,解析技术瓶颈与未来方向

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation

  • 结合生成AI与分治策略,提升长视频生成可扩展性
  • 现有模型如Sora最长仅支持1分钟视频生成
  • 适合关注视频生成、多模态模型的研究者参考

一张图像可传递千言万语,而由数百甚至数千帧图像组成的视频则能讲述更复杂的故事。尽管多模态大语言模型(MLLM)取得显著进展,长视频生成仍是重大挑战。截至本文撰写时,当前最先进的系统OpenAI Sora仍仅支持生成时长不超过一分钟的视频。这一局限源于长视频生成的复杂性,其不仅需要生成式AI对密度函数进行近似,还涉及规划、故事发展以及时空一致性维护等关键难题。将生成式AI与分治策略相结合,或可提升长视频生成的可扩展性并增强控制力。本文综述了长视频生成的当前研究格局,涵盖基础技术如GANs和扩散模型、视频生成策略、大规模训练数据集、长视频质量评估指标,以及未来研究方向。我们认为该综述可为长视频生成领域的未来发展提供全面基础信息。

原文摘要 · Abstract (English)

An image may convey a thousand words, but a video composed of hundreds or thousands of image frames tells a more intricate story. Despite significant progress in multimodal large language models (MLLMs), generating extended videos remains a formidable challenge. As of this writing, OpenAI's Sora, the current state-of-the-art system, is still limited to producing videos that are up to one minute in length. This limitation stems from the complexity of long video generation, which requires more than generative AI techniques for approximating density functions essential aspects such as planning, story development, and maintaining spatial and temporal consistency present additional hurdles. Integrating generative AI with a divide-and-conquer approach could improve scalability for longer videos while offering greater control. In this survey, we examine the current landscape of long video generation, covering foundational techniques like GANs and diffusion models, video generation strategies, large-scale training datasets, quality metrics for evaluating long videos, and future research areas to address the limitations of the existing video generation capabilities. We believe it would serve as a comprehensive foundation, offering extensive information to guide future advancements and research in the field of long video generation.

视频生成多模态大模型综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。