通过可控相机实现大范围动态场景生成,突破视频生成视角局限。
CameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models
- 分步扩展动态场景生成:先增强单片段动态性,再实现多视角连贯合成。
- 支持大幅相机移动,生成视频的视角覆盖范围显著超越以往方法。
- 适合需要自由相机轨迹控制的虚拟拍摄、游戏开发等应用。
本文提出CameraCtrl II框架,通过可控相机实现大规模动态场景生成。现有相机条件视频生成模型在大幅相机运动时易丢失动态效果且视角范围受限。本工作采用渐进式策略:首先构建包含丰富动态与相机参数标注的大规模数据集,设计轻量级相机注入模块及训练方案,以保留预训练模型的动态特性;在此基础上,支持用户迭代指定相机轨迹,生成连贯的长视频序列。实验表明,CameraCtrl II在多种场景下均能实现比先前方法更广的空间探索能力。
原文摘要 · Abstract (English)
This paper introduces CameraCtrl II, a framework that enables large-scale dynamic scene exploration through a camera-controlled video diffusion model. Previous camera-conditioned video generative models suffer from diminished video dynamics and limited range of viewpoints when generating videos with large camera movement. We take an approach that progressively expands the generation of dynamic scenes -- first enhancing dynamic content within individual video clip, then extending this capability to create seamless explorations across broad viewpoint ranges. Specifically, we construct a dataset featuring a large degree of dynamics with camera parameter annotations for training while designing a lightweight camera injection module and training scheme to preserve dynamics of the pretrained models. Building on these improved single-clip techniques, we enable extended scene exploration by allowing users to iteratively specify camera trajectories for generating coherent video sequences. Experiments across diverse scenarios demonstrate that CameraCtrl Ii enables camera-controlled dynamic scene synthesis with substantially wider spatial exploration than previous approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。