从单张图片生成连贯3D场景,提升重建精度与一致性。
Coherent 3D Scene Diffusion From a Single RGB Image
- 用扩散模型联合优化场景中所有物体的位姿和几何
- 在SUN RGB-D上AP3D提升12.04%,Pix3D上F-Score提升13.43%
- 无需完整标注也能训练,适合缺乏真实标签的数据集
我们提出一种基于扩散的新方法,从单张RGB图像实现连贯的3D场景重建。该方法利用图像条件化的3D场景扩散模型,同时对场景中所有物体的位姿和几何进行去噪。受任务固有不确定性启发,通过同时条件化所有场景物体来学习生成式场景先验,捕捉场景上下文并让模型在扩散过程中学习物体间关系。此外,我们设计了一种高效的表面对齐损失,在缺乏完整真值标注(常见于公开数据集)时仍能有效训练,该损失使用表达力强的形状表示,可直接从中间形状预测采样点。将单张RGB图像3D场景重建建模为条件扩散过程后,本方法超越现有最先进方法,在SUN RGB-D上实现AP3D提升12.04%,在Pix3D上F-Score提升13.43%。
原文摘要 · Abstract (English)
We present a novel diffusion-based approach for coherent 3D scene reconstruction from a single RGB image. Our method utilizes an image-conditioned 3D scene diffusion model to simultaneously denoise the 3D poses and geometries of all objects within the scene. Motivated by the ill-posed nature of the task and to obtain consistent scene reconstruction results, we learn a generative scene prior by conditioning on all scene objects simultaneously to capture the scene context and by allowing the model to learn inter-object relationships throughout the diffusion process. We further propose an efficient surface alignment loss to facilitate training even in the absence of full ground-truth annotation, which is common in publicly available datasets. This loss leverages an expressive shape representation, which enables direct point sampling from intermediate shape predictions. By framing the task of single RGB image 3D scene reconstruction as a conditional diffusion process, our approach surpasses current state-of-the-art methods, achieving a 12.04% improvement in AP3D on SUN RGB-D and a 13.43% increase in F-Score on Pix3D.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。