arXiv:2601.16214cs.CV2026-01被引 2

用3D高斯解码提升视频扩散模型的摄像机控制精度

CamPilot: Improving Camera Control in Video Diffusion Model with Efficient Camera Reward Feedback

  • 通过3D高斯解码融合相机位姿,实现像素级对齐奖励
  • 在RealEstate10K和WorldScore上显著提升相机控制效果
  • 适合关注视频生成中相机精确控制的研究者

近期基于摄像机控制的视频扩散模型虽提升了视频与摄像机的一致性,但控制能力仍受限。现有奖励反馈学习方法面临三大挑战:现有奖励模型无法评估视频-摄像机对齐;解码潜空间为RGB视频带来巨大计算开销;视频解码中常忽略3D几何信息。为此,我们提出高效摄像机感知的3D解码器,将视频潜变量与相机位姿共同解码为3D高斯表示。在此过程中,相机位姿既作为输入,也作为投影参数。若视频潜变量与相机位姿不匹配,将导致3D结构几何失真,进而产生模糊渲染结果。基于此特性,我们以新视角渲染图与真实图像间的像素一致性作为奖励信号。为应对随机性,进一步引入可见性项,仅对几何映射得到的确定区域进行监督。在RealEstate10K和WorldScore基准上的大量实验验证了方法的有效性。

原文摘要 · Abstract (English)

Recent advances in camera-controlled video diffusion models have significantly improved video-camera alignment. However, the camera controllability still remains limited. In this work, we build upon Reward Feedback Learning and aim to further improve camera controllability. However, directly borrowing existing ReFL approaches faces several challenges. First, current reward models lack the capacity to assess video-camera alignment. Second, decoding latent into RGB videos for reward computation introduces substantial computational overhead. Third, 3D geometric information is typically neglected during video decoding. To address these limitations, we introduce an efficient camera-aware 3D decoder that decodes video latent into 3D representations for reward quantization. Specifically, video latent along with the camera pose are decoded into 3D Gaussians. In this process, the camera pose not only acts as input, but also serves as a projection parameter. Misalignment between the video latent and camera pose will cause geometric distortions in the 3D structure, resulting in blurry renderings. Based on this property, we explicitly optimize pixel-level consistency between the rendered novel views and ground-truth ones as reward. To accommodate the stochastic nature, we further introduce a visibility term that selectively supervises only deterministic regions derived via geometric warping. Extensive experiments conducted on RealEstate10K and WorldScore benchmarks demonstrate the effectiveness of our proposed method. Project page: \href{https://a-bigbao.github.io/CamPilot/}{CamPilot Page}.

视频生成扩散模型相机控制3D高斯

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。