单张图像重建完整3D场景,填补遮挡区域缺失结构
VolFill: Single-View Amodal 3D Scene Reconstruction with Volumetric Flow Matching

- 用体积流匹配生成完整3D结构,不依赖像素对齐回归
- 在SCRREAM和NRGB-D数据集上显著优于现有基线方法
- 支持大规模表面提取与占据查询,适合需要完整几何的场景理解
从单张RGB图像重建完整场景几何仍具挑战性,尤其在视觉证据不完整时需推断隐藏结构。我们提出VolFill,一种生成式框架,直接预测完整场景的3D结构,而非传统像素对齐回归。该方法采用混合3D VAE将稀疏截断无符号距离函数网格压缩至紧凑潜在空间,并结合潜在扩散变压器对表示进行去噪以恢复完整场景。生成过程基于几何基础模型进行条件约束,利用丰富的空间先验实现鲁棒推理。与受限于每射线约束或非结构化点云查询的方法不同,VolFill提供结构化表示,支持大规模直接表面提取与占据查询。在SCRREAM和NRGB-D数据集上的大量实验表明,本方法显著超越现有基线,为整体空间理解提供坚实基础。
原文摘要 · Abstract (English)
Reconstructing the complete geometry of a scene from a single RGB image remains challenging - especially when inferring hidden structures where visual evidence is incomplete. We introduce VolFill, a generative framework that predicts the 3D structure of the complete scene rather than relying on traditional pixel-aligned regression. Our method utilizes a hybrid 3D VAE to compress sparse truncated unsigned distance function grids into a compact latent space, paired with a latent Diffusion Transformer that denoises this representation to recover the complete scene. We condition the generation on geometry foundation models, leveraging rich spatial priors for robust reasoning. Unlike existing methods limited by per-ray constraints or unstructured point-cloud queries, VolFill provides a structured representation that supports direct surface extraction and occupancy queries at scale. Extensive experiments on the SCRREAM and NRGB-D datasets demonstrate that our approach significantly outperforms current baselines, providing a robust foundation for holistic spatial understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。