用隐式3D场景表示生成电影级视频,实现动态主体与镜头运动的精准控制。
CineScene: Implicit 3D as Effective Scene Representation for Cinematic Video Generation
- 通过VGGT编码场景图,隐式注入3D空间先验到文生视频模型中
- 在未见环境中实现大范围镜头移动下的场景一致性,视频质量达当前最优
- 适合影视制作、虚拟拍摄等需要高效可控场景生成的领域
电影视频制作需精确控制场景-主体构图与镜头运动,但实拍成本高昂,因需搭建物理布景。为此,我们提出解耦场景上下文的电影级视频生成任务:给定静态环境的多张图像,目标是合成高质量视频,包含动态主体并保持场景一致性和用户指定的相机轨迹。我们提出CineScene框架,利用隐式3D感知场景表征进行视频生成。核心创新在于一种新颖的上下文条件机制,通过VGGT将场景图像编码为视觉表征,并以附加上下文拼接方式隐式注入3D先验,使预训练文生视频模型能生成受相机控制、场景一致的视频。为提升鲁棒性,训练时引入简单的输入场景图像随机打乱策略。针对训练数据不足问题,我们基于Unreal Engine 5构建了一个解耦场景数据集,包含带与不带动态主体的配对视频、全景静态场景图及对应相机轨迹。实验表明,CineScene在场景一致性电影视频生成上达到当前最优性能,可处理大范围镜头运动,并在多样化环境中展现良好泛化能力。
原文摘要 · Abstract (English)
Cinematic video production requires control over scene-subject composition and camera movement, but live-action shooting remains costly due to the need for constructing physical sets. To address this, we introduce the task of cinematic video generation with decoupled scene context: given multiple images of a static environment, the goal is to synthesize high-quality videos featuring dynamic subject while preserving the underlying scene consistency and following a user-specified camera trajectory. We present CineScene, a framework that leverages implicit 3D-aware scene representation for cinematic video generation. Our key innovation is a novel context conditioning mechanism that injects 3D-aware features in an implicit way: By encoding scene images into visual representations through VGGT, CineScene injects spatial priors into a pretrained text-to-video generation model by additional context concatenation, enabling camera-controlled video synthesis with consistent scenes and dynamic subjects. To further enhance the model's robustness, we introduce a simple yet effective random-shuffling strategy for the input scene images during training. To address the lack of training data, we construct a scene-decoupled dataset with Unreal Engine 5, containing paired videos of scenes with and without dynamic subjects, panoramic images representing the underlying static scene, along with their camera trajectories. Experiments show that CineScene achieves state-of-the-art performance in scene-consistent cinematic video generation, handling large camera movements and demonstrating generalization across diverse environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。