arXiv:2411.15262cs.CV2024-11CVPR被引 38

构建首个面向长视频生成的分层电影级数据集

MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation

  • 按电影级结构组织多场景叙事数据
  • 保证角色外观与声音跨场景一致
  • 适合研究长视频生成与一致性建模

近期视频生成模型(如 Stable Video Diffusion)在短时单场景视频生成上表现良好,但难以生成包含多场景、连贯叙事和角色一致性的长视频。目前缺乏专门用于长视频生成模型训练与评估的公开数据集。本文提出 MovieBench:一个面向长视频生成的分层电影级数据集,具备三大特性:(1) 包含丰富连贯剧情的电影级长视频;(2) 跨场景的角色外观与音频一致性;(3) 分层数据结构,包含高层电影信息与细粒度镜头描述。实验表明,该数据集揭示了跨多场景保持角色身份一致性的新挑战。数据集将公开并持续维护,推动长视频生成技术发展。数据获取地址:https://weijiawu.github.io/MovieBench/

原文摘要 · Abstract (English)

Recent advancements in video generation models, like Stable Video Diffusion, show promising results, but primarily focus on short, single-scene videos. These models struggle with generating long videos that involve multiple scenes, coherent narratives, and consistent characters. Furthermore, there is no publicly available dataset tailored for the analysis, evaluation, and training of long video generation models. In this paper, we present MovieBench: A Hierarchical Movie-Level Dataset for Long Video Generation, which addresses these challenges by providing unique contributions: (1) movie-length videos featuring rich, coherent storylines and multi-scene narratives, (2) consistency of character appearance and audio across scenes, and (3) hierarchical data structure contains high-level movie information and detailed shot-level descriptions. Experiments demonstrate that MovieBench brings some new insights and challenges, such as maintaining character ID consistency across multiple scenes for various characters. The dataset will be public and continuously maintained, aiming to advance the field of long video generation. Data can be found at: https://weijiawu.github.io/MovieBench/.

长视频生成分层数据集角色一致性电影级数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。