联合生成带精准相机控制的RGB-D视频,几何一致性更强。
IDCNet: Guided Video Diffusion for Metric-Consistent RGBD Scene Generation with Precise Camera Control
- 用统一的几何感知扩散模型同时生成图像与深度图。
- 在真实相机轨迹下生成序列,帧间几何一致性显著提升。
- 适合需要高精度3D重建的虚拟场景生成任务。
我们提出IDC-Net(图像-深度一致性网络),一种在显式相机轨迹控制下生成RGB-D视频序列的新框架。不同于将RGB与深度分别生成的方法,IDC-Net在统一的几何感知扩散模型中联合合成图像与对应深度图,强化了帧间的空间与几何对齐,从而实现更精确的相机控制。为支持该相机条件模型的训练并确保高几何保真度,我们构建了一个具有度量对齐的RGB视频、深度图与精确相机位姿的相机-图像-深度一致数据集,提供显著提升的帧间几何一致性监督。此外,我们引入几何感知注意力模块,实现细粒度相机控制。大量实验表明,IDC-Net在生成序列的视觉质量与几何一致性上均优于现有方法。值得注意的是,生成的RGB-D序列可直接用于下游3D场景重建任务,无需额外后处理,体现了联合学习框架的实际优势。
原文摘要 · Abstract (English)
We present IDC-Net (Image-Depth Consistency Network), a novel framework designed to generate RGB-D video sequences under explicit camera trajectory control. Unlike approaches that treat RGB and depth generation separately, IDC-Net jointly synthesizes both RGB images and corresponding depth maps within a unified geometry-aware diffusion model. The joint learning framework strengthens spatial and geometric alignment across frames, enabling more precise camera control in the generated sequences. To support the training of this camera-conditioned model and ensure high geometric fidelity, we construct a camera-image-depth consistent dataset with metric-aligned RGB videos, depth maps, and accurate camera poses, which provides precise geometric supervision with notably improved inter-frame geometric consistency. Moreover, we introduce a geometry-aware transformer block that enables fine-grained camera control, enhancing control over the generated sequences. Extensive experiments show that IDC-Net achieves improvements over state-of-the-art approaches in both visual quality and geometric consistency of generated scene sequences. Notably, the generated RGB-D sequences can be directly feed for downstream 3D Scene reconstruction tasks without extra post-processing steps, showcasing the practical benefits of our joint learning framework. See more at https://idcnet-scene.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。