arXiv:2501.06006cs.CV2025-01被引 22

仅用一张图和相机路径生成3D飞掠视频,细节保真度高。

CamCtrl3D: Single-Image Scene Exploration with Precise 3D Camera Control

  • 通过四类条件融合控制相机轨迹生成视频
  • 在多个数据集上实现当前最优的视图一致性与细节保留
  • 适合需要高质量3D场景探索的应用开发者

我们提出一种从单张图像和给定相机轨迹生成场景飞掠视频的方法。基于图像到视频的潜在扩散模型,通过四种技术将相机轨迹条件引入其UNet去噪器:(1) 在时间模块中直接使用原始相机外参;(2) 使用包含相机射线和方向的图像;(3) 将初始图像重投影至后续帧作为条件;(4) 利用2D<=>3D变换器引入全局3D表示以隐式编码相机位姿。所有条件通过类似ControlNet的架构融合。我们设计了一项评估指标,用于衡量视频整体质量与视角变化下的细节保持能力,分析各条件的权衡并识别最优组合。我们在数据集中校准相机位置以保证尺度一致性,并训练出名为CamCtrl3D的场景探索模型,实现在多个基准上的最先进性能。

原文摘要 · Abstract (English)

We propose a method for generating fly-through videos of a scene, from a single image and a given camera trajectory. We build upon an image-to-video latent diffusion model. We condition its UNet denoiser on the camera trajectory, using four techniques. (1) We condition the UNet's temporal blocks on raw camera extrinsics, similar to MotionCtrl. (2) We use images containing camera rays and directions, similar to CameraCtrl. (3) We reproject the initial image to subsequent frames and use the resulting video as a condition. (4) We use 2D<=>3D transformers to introduce a global 3D representation, which implicitly conditions on the camera poses. We combine all conditions in a ContolNet-style architecture. We then propose a metric that evaluates overall video quality and the ability to preserve details with view changes, which we use to analyze the trade-offs of individual and combined conditions. Finally, we identify an optimal combination of conditions. We calibrate camera positions in our datasets for scale consistency across scenes, and we train our scene exploration model, CamCtrl3D, demonstrating state-of-theart results.

3D生成相机控制扩散模型视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。