让视频生成可自由安排镜头,还能控制角色动作和场景。
MultiShotMaster: A Controllable Multi-Shot Video Generation Framework
- 用两种新位置编码实现镜头切换与时空定位控制
- 支持灵活镜头数量与时长,保持叙事连贯性
- 适合需要精细控制视频镜头的创作者使用
当前视频生成技术擅长单镜头片段,但在生成具有灵活镜头安排、连贯叙事的多镜头视频方面仍存在困难。为此,我们提出MultiShotMaster框架,实现高度可控的多镜头视频生成。通过扩展预训练单镜头模型,引入两种新型RoPE:多镜头叙事RoPE在镜头切换处施加显式相位偏移,支持灵活镜头排列并保持时间叙事顺序;时空位置感知RoPE融合参考标记与定位信号,实现时空锚定的参考注入。此外,为克服数据稀缺问题,构建自动化数据标注流程,提取多镜头视频、字幕、跨镜头定位信号及参考图像。该框架利用内在架构特性,支持文本驱动的镜头间一致性、主体运动自定义以及背景引导的场景定制,镜头数量与持续时间均可灵活配置。大量实验表明,本框架在性能与可控性上均表现卓越。
原文摘要 · Abstract (English)
Current video generation techniques excel at single-shot clips but struggle to produce narrative multi-shot videos, which require flexible shot arrangement, coherent narrative, and controllability beyond text prompts. To tackle these challenges, we propose MultiShotMaster, a framework for highly controllable multi-shot video generation. We extend a pretrained single-shot model by integrating two novel variants of RoPE. First, we introduce Multi-Shot Narrative RoPE, which applies explicit phase shift at shot transitions, enabling flexible shot arrangement while preserving the temporal narrative order. Second, we design Spatiotemporal Position-Aware RoPE to incorporate reference tokens and grounding signals, enabling spatiotemporal-grounded reference injection. In addition, to overcome data scarcity, we establish an automated data annotation pipeline to extract multi-shot videos, captions, cross-shot grounding signals and reference images. Our framework leverages the intrinsic architectural properties to support multi-shot video generation, featuring text-driven inter-shot consistency, customized subject with motion control, and background-driven customized scene. Both shot count and duration are flexibly configurable. Extensive experiments demonstrate the superior performance and outstanding controllability of our framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。