arXiv:2511.23127cs.CV2025-11被引 9

DualCamCtrl通过双分支扩散模型实现几何感知的相机控制视频生成。

DualCamCtrl: Dual-Branch Diffusion Model for Geometry-Aware Camera-Controlled Video Generation

  • 双分支结构同步生成一致的RGB与深度序列。
  • 相比之前方法,相机运动误差降低40%以上。
  • 适合需要精确相机轨迹控制的视频生成场景。

本文提出DualCamCtrl,一种端到端的扩散模型,用于相机控制的视频生成。现有方法虽将相机姿态表示为基于射线的条件,但常缺乏足够的场景理解与几何感知。DualCamCtrl通过双分支框架,相互生成一致的RGB与深度序列,并引入语义引导的互对齐机制(SIGMA),在语义引导下实现RGB与深度的双向融合。该设计有效解耦外观与几何建模,使生成视频更忠实于指定相机轨迹。我们还分析了深度与相机姿态在去噪不同阶段的影响,发现早期与晚期阶段分别负责全局结构形成与局部细节优化。大量实验表明,DualCamCtrl在相机控制视频生成上表现更优,相机运动误差较先前方法降低超过40%。

原文摘要 · Abstract (English)

This paper presents DualCamCtrl, a novel end-to-end diffusion model for camera-controlled video generation. Recent works have advanced this field by representing camera poses as ray-based conditions, yet they often lack sufficient scene understanding and geometric awareness. DualCamCtrl specifically targets this limitation by introducing a dual-branch framework that mutually generates camera-consistent RGB and depth sequences. To harmonize these two modalities, we further propose the Semantic Guided Mutual Alignment (SIGMA) mechanism, which performs RGB-depth fusion in a semantics-guided and mutually reinforced manner. These designs collectively enable DualCamCtrl to better disentangle appearance and geometry modeling, generating videos that more faithfully adhere to the specified camera trajectories. Additionally, we analyze and reveal the distinct influence of depth and camera poses across denoising stages and further demonstrate that early and late stages play complementary roles in forming global structure and refining local details. Extensive experiments demonstrate that DualCamCtrl achieves more consistent camera-controlled video generation, with over 40\% reduction in camera motion errors compared with prior methods. Our project page: https://soyouthinkyoucantell.github.io/dualcamctrl-page/

视频生成扩散模型相机控制几何感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。