用分步生成法创作连贯短片,从剧本到画面一气呵成。
Captain Cinema: Towards Short Movie Generation
- 先生成关键帧规划剧情,再合成中间动态视频
- 在多场景长叙事下保持画面与剧情一致
- 适合影视自动化创作,效率高且质量优
我们提出 Captain Cinema,一个面向短片生成的框架。给定详细文本描述的电影剧情,该方法首先生成一组关键帧以规划完整叙事,确保故事与视觉呈现(如场景和角色)的长期一致性,此为自上而下的关键帧规划。随后,这些关键帧作为条件信号输入视频合成模型,该模型支持长上下文学习,生成关键帧间的时空动态,称为自下而上的视频合成。为稳定高效生成多场景长叙事影视作品,我们设计了针对长上下文视频数据的多模态扩散变换器(MM-DiT)交错训练策略,并在专门构建的交错数据对电影数据集上进行训练。实验表明,Captain Cinema 在自动创建高质量、视觉连贯且叙事一致的短片方面表现优异,兼具高效性与生成稳定性。
原文摘要 · Abstract (English)
We present Captain Cinema, a generation framework for short movie generation. Given a detailed textual description of a movie storyline, our approach firstly generates a sequence of keyframes that outline the entire narrative, which ensures long-range coherence in both the storyline and visual appearance (e.g., scenes and characters). We refer to this step as top-down keyframe planning. These keyframes then serve as conditioning signals for a video synthesis model, which supports long context learning, to produce the spatio-temporal dynamics between them. This step is referred to as bottom-up video synthesis. To support stable and efficient generation of multi-scene long narrative cinematic works, we introduce an interleaved training strategy for Multimodal Diffusion Transformers (MM-DiT), specifically adapted for long-context video data. Our model is trained on a specially curated cinematic dataset consisting of interleaved data pairs. Our experiments demonstrate that Captain Cinema performs favorably in the automated creation of visually coherent and narrative consistent short movies in high quality and efficiency. Project page: https://thecinema.ai
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。