arXiv:2601.10214cs.CVcs.GR2026-01被引 2

通过深度图引导重渲染,实现精准相机控制的视频生成。

Beyond Inpainting: Unleash 3D Understanding for Precise Camera-Controlled Video Generation

  • 用源视频和目标视角的深度图作为双重条件输入
  • 在8000个动态场景上实现高质量相机轨迹控制
  • 适合需要精确相机编辑的影视与游戏应用

相机控制在条件视频生成中已广泛研究,但精确调整相机轨迹并忠实保留视频内容仍具挑战。主流方法通过变形3D表示来实现相机控制,却未能充分挖掘视频扩散模型(VDMs)的3D先验,常陷入补绘陷阱,导致主体不一致与生成质量下降。为此,我们提出DepthDirector——一种具备精准相机可控性的视频重渲染框架。该方法利用显式3D表示生成的深度视频作为相机控制引导,可在新相机轨迹下忠实还原输入视频的动态场景。具体地,设计视点-内容双流条件机制,将源视频与目标视角下扭曲的深度序列共同注入预训练视频生成模型。这种几何引导信号使VDMs理解相机运动,激发其3D理解能力,从而实现精准控制与内容一致性。此外,采用轻量级LoRA基视频扩散适配器进行训练,完整保留VDMs的知识先验。我们还基于虚幻引擎5构建了大规模多相机同步数据集MultiCam-WarpData,包含1000个动态场景中的8000个视频。大量实验表明,DepthDirector在相机可控性与视觉质量上均优于现有方法。代码与数据集将公开。

原文摘要 · Abstract (English)

Camera control has been extensively studied in conditioned video generation; however, performing precisely altering the camera trajectories while faithfully preserving the video content remains a challenging task. The mainstream approach to achieving precise camera control is warping a 3D representation according to the target trajectory. However, such methods fail to fully leverage the 3D priors of video diffusion models (VDMs) and often fall into the Inpainting Trap, resulting in subject inconsistency and degraded generation quality. To address this problem, we propose DepthDirector, a video re-rendering framework with precise camera controllability. By leveraging the depth video from explicit 3D representation as camera-control guidance, our method can faithfully reproduce the dynamic scene of an input video under novel camera trajectories. Specifically, we design a View-Content Dual-Stream Condition mechanism that injects both the source video and the warped depth sequence rendered under the target viewpoint into the pretrained video generation model. This geometric guidance signal enables VDMs to comprehend camera movements and leverage their 3D understanding capabilities, thereby facilitating precise camera control and consistent content generation. Next, we introduce a lightweight LoRA-based video diffusion adapter to train our framework, fully preserving the knowledge priors of VDMs. Additionally, we construct a large-scale multi-camera synchronized dataset named MultiCam-WarpData using Unreal Engine 5, containing 8K videos across 1K dynamic scenes. Extensive experiments show that DepthDirector outperforms existing methods in both camera controllability and visual quality. Our code and dataset will be publicly available.

视频生成相机控制3D理解扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。