arXiv:2507.07202cs.CV2025-07ICCV综述被引 7

综述长视频叙事生成的架构与一致性技术,解决多角色连贯性难题。

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality

  • 梳理32篇论文,归纳生成长视频的关键架构与训练策略
  • 揭示当前方法在150秒以上视频中存在帧冗余和时序多样性低的问题
  • 提出新分类体系,适合研究长视频生成与影视级质量优化的学者

尽管视频生成模型取得了显著进展,现有最先进方法生成的视频通常仅持续5-16秒,常被称作“长视频”。超过16秒的视频难以维持角色形象与场景布局的一致性,尤其是多主体长视频仍无法保证角色一致性和动作连贯性。虽有部分方法可生成长达150秒的视频,但常伴随帧冗余和时间维度多样性不足的问题。近期工作尝试生成包含多个角色、叙事连贯且细节高保真的长视频。我们系统分析了32篇视频生成相关论文,识别出能稳定实现这些特性的关键架构组件与训练策略。同时构建了一个全新的方法分类体系,并提供对比表格,按架构设计与性能特征对论文进行归类。

原文摘要 · Abstract (English)

Despite the significant progress that has been made in video generative models, existing state-of-the-art methods can only produce videos lasting 5-16 seconds, often labeled "long-form videos". Furthermore, videos exceeding 16 seconds struggle to maintain consistent character appearances and scene layouts throughout the narrative. In particular, multi-subject long videos still fail to preserve character consistency and motion coherence. While some methods can generate videos up to 150 seconds long, they often suffer from frame redundancy and low temporal diversity. Recent work has attempted to produce long-form videos featuring multiple characters, narrative coherence, and high-fidelity detail. We comprehensively studied 32 papers on video generation to identify key architectural components and training strategies that consistently yield these qualities. We also construct a comprehensive novel taxonomy of existing methods and present comparative tables that categorize papers by their architectural designs and performance characteristics.

视频生成长视频一致性叙事生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。