用可学习的镜头标记实现电影级多镜头视频生成
ShotPlan: Cinematic Video Generation with Learnable Planning Token

- 引入可学习的拍摄规划标记,控制镜头切换时机
- 在多个数据集上显著提升镜头间一致性与叙事连贯性
- 适合需要精细镜头调度的影视生成场景
当前视频生成模型在单镜头生成上表现优异,但在需要连贯叙事和有效多镜头组合的电影级视频生成上仍受限。为此,我们提出ShotPlan,一种基于视频扩散模型的显式多镜头电影视频生成框架。该方法引入可学习的规划标记,捕捉镜头级别的过渡线索,并可无缝集成到原始视频生成标记中以控制切换时间点。不同于标准视频生成标记,所提出的规划标记采用分数时间旋转位置编码(FRoPE),实现帧级的镜头过渡建模。实验表明,ShotPlan显著优于现有电影视频生成方法,在镜头管理灵活性和镜头间一致性方面表现更优。
原文摘要 · Abstract (English)
Current video generation models achieve impressive results in single-shot generation, yet remain limited in cinematic video generation, where coherent narratives and effective multi-shot composition require explicit shot planning. To address this challenge, we propose ShotPlan, a framework for explicit multi-shot cinematic video generation built upon a video diffusion foundation model. Our method introduces learnable planning tokens that capture shot-level transition cues and can be seamlessly integrated with the original video generation tokens to control transition timestamps. Unlike standard video generation tokens, the proposed planning tokens are equipped with Fractional Temporal Rotary Position Embedding (FRoPE), enabling shot transitions to be modeled at the frame level. Experiments demonstrate that ShotPlan significantly outperforms existing cinematic video generation methods, offering more flexible shot management and stronger inter-shot consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。