用生成模型从稀疏视角重建复杂多物体场景,效果远超此前方法。
Reconstruction by Generation: 3D Multi-Object Scene Reconstruction from Sparse Observations

- 基于合成场景和3D形状先验,联合生成物体形状与位姿。
- 在严重遮挡下仍能提升几何、纹理和位姿重建准确率。
- 训练数据量少80%却性能更优,适合机器人仿真应用。
从稀疏观测中精确重建复杂的完整多物体场景,仍是计算机视觉的核心挑战,也是实现机器人可扩展可靠仿真的关键步骤。本文提出RecGen,一种生成式框架,用于在遮挡和部分可见条件下,从单张或多张RGB-D图像中进行物体与部件形状及位姿的联合概率估计。通过利用组合式合成场景生成与强3D形状先验,RecGen能泛化于多种物体类型与真实环境。在复杂且严重遮挡的数据集上,RecGen达到当前最优性能,稳健处理严重遮挡、对称物体、部件结构以及复杂几何与纹理。尽管训练所用网格数比此前最优方法SAM3D少近80%,其在几何形状质量上仍领先30.1%,纹理重建提升9.1%,位姿估计提高33.9%。
原文摘要 · Abstract (English)
Accurately reconstructing complex full multi-object scenes from sparse observations remains a core challenge in computer vision and a key step toward scalable and reliable simulation for robotics. In this work, we introduce RecGen, a generative framework for probabilistic joint estimation of object and part shapes, as well as their pose under occlusion and partial visibility from one or multiple RGB-D images. By leveraging compositional synthetic scene generation and strong 3D shape priors, RecGen generalizes across diverse object types and real-world environments. RecGen achieves state-of-the-art performance on complex, heavily occluded datasets, robustly handling severe occlusions, symmetric objects, object parts, and intricate geometry and texture. Despite using nearly 80% fewer training meshes than the previous state of the art SAM3D, RecGen outperforms it by 30.1% in geometric shape quality, 9.1% in texture reconstruction, and 33.9% in pose estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。