统一控制人物、事件、镜头与转场,生成电影级视频
CineOrchestra: Unified Entity-Centric Conditioning for Cinematic Video Generation

- 用统一实体中心条件框架,同时管理人物、事件、镜头和转场
- 在两个新基准上超越六种单项专精模型,转场时间更准
- 适合影视创作、视频编辑等需要精细控制的场景
电影视频包含多个角色在特定时刻的动作或互动,配合精心设计的镜头运动,并通过镜头切换串联。这些元素要求细粒度控制能力,远超现有文本到视频模型的能力。当前研究仅分别处理多角色个性化、时间控制、多镜头合成或镜头运动;尚无框架能联合建模这四项。我们提出CineOrchestra,一个统一的视频扩散模型,可同时控制角色、事件、镜头和镜头转换。核心洞察是:这些异构电影元素共享基本结构——每个都是在特定时间段内活动的实体,均可通过统一的实体中心条件原语表达,辅以参考图像提供视觉实体信息。该形式将架构挑战简化为单一位置编码问题,我们通过两种无需参数的协同旋转编码解决:(a) 区间采样的时序RoPE,实现不同持续时间事件间一致的关注行为;(b) 二维实体-时序交叉注意力RoPE,明确区分各实体条件并路由至对应时空区域。在两个新基准上,CineOrchestra在密集描述跟随和镜头转换时机任务中优于六种单项专家模型,在配对用户研究和组件消融中均保持稳定提升。
原文摘要 · Abstract (English)
Cinematic video depicts multiple subjects acting or interacting at specific moments, captured with deliberate camera movement, and stitched together by shot transitions. Together, these elements demand a level of fine-grained control beyond current text-to-video models. Existing work addresses each axis in isolation: multi-subject personalization, temporal control, multi-shot synthesis, or camera control; no prior framework jointly integrates all four. We present CineOrchestra, a unified video diffusion model that controls subjects, events, cameras, and shot transitions simultaneously. Our key insight is that these heterogeneous cinematic elements share a fundamental structure: each is an entity acting over a specific temporal interval, which can therefore all be expressed through one shared structure of entity-centric conditioning primitives, augmented with reference images for visual entities. This formulation reduces the architectural challenge to a single positional encoding problem, which we solve with two parameter-free coordinated rotary embeddings: (a) an interval-sampled temporal RoPE that yields consistent attention behavior across events of dramatically varying duration, and (b) a 2D entity-temporal cross-attention RoPE that disambiguates per-entity conditions and routes each to its corresponding spatiotemporal region. On two new benchmarks, CineOrchestra outperforms six per-axis specialists on dense caption following and shot-transition timing, with consistent gains in a pairwise user study and component ablations. Project page: https://snap-research.github.io/CineOrchestra
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。