让AI一次生成连贯长篇电影叙事,支持角色与镜头的全局一致性。
HoloCine: Holistic Generation of Cinematic Multi-Shot Long Video Narratives
- 通过窗口交叉注意力实现文本提示精准定位到特定镜头。
- 采用稀疏跨镜头自注意力,支持分钟级视频高效生成。
- 具备角色记忆与电影运镜直觉,适合影视创作与交互设计。
当前最先进的文生视频模型擅长生成孤立片段,但在构建连贯的多镜头叙事方面表现不足,而这是讲故事的核心。我们提出HoloCine,通过整体化生成整个场景,确保从第一镜到最后一镜的全局一致性,弥合了这一‘叙事鸿沟’。其架构利用窗口交叉注意力机制,将文本提示精确定位至特定镜头;同时采用稀疏跨镜头自注意力模式(镜头内密集、镜头间稀疏),保障分钟级生成的效率。除在叙事连贯性上达到新基准外,HoloCine还展现出显著的涌现能力:对角色和场景具有持续记忆,以及对电影手法的直观理解。本工作标志着从片段合成迈向自动化电影制作的转折点,使端到端的影视创作成为可能。代码已开源:https://holo-cine.github.io/。
原文摘要 · Abstract (English)
State-of-the-art text-to-video models excel at generating isolated clips but fall short of creating the coherent, multi-shot narratives, which are the essence of storytelling. We bridge this "narrative gap" with HoloCine, a model that generates entire scenes holistically to ensure global consistency from the first shot to the last. Our architecture achieves precise directorial control through a Window Cross-Attention mechanism that localizes text prompts to specific shots, while a Sparse Inter-Shot Self-Attention pattern (dense within shots but sparse between them) ensures the efficiency required for minute-scale generation. Beyond setting a new state-of-the-art in narrative coherence, HoloCine develops remarkable emergent abilities: a persistent memory for characters and scenes, and an intuitive grasp of cinematic techniques. Our work marks a pivotal shift from clip synthesis towards automated filmmaking, making end-to-end cinematic creation a tangible future. Our code is available at: https://holo-cine.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。