利用重复物体提升单视角3D场景重建精度
FurnSet: Exploiting Repeats for 3D Scene Reconstruction

- 通过对象级CLSToken和集合感知注意力,识别并聚合重复物体
- 在3D-Future和3D-Front上实现更优的场景重建质量
- 适合需要高精度3D场景重建的室内设计与机器人应用
单视图3D场景重建旨在推断物体几何与空间布局。现有方法通常独立重建物体或依赖隐式场景上下文,未能利用真实场景中常见的重复实例。我们提出FurnSet框架,显式识别并利用重复物体实例以提升重建效果。该方法引入每个物体的CLS token和集合感知自注意力机制,将相同实例分组,并聚合其互补观测信息,实现联合重建。进一步结合场景级与物体级条件引导物体重建,随后利用物体点云的3D与2D投影损失进行布局优化,实现场景对齐。在3D-Future和3D-Front数据集上的实验表明,利用重复性可显著提升3D场景重建质量。
原文摘要 · Abstract (English)
Single-view 3D scene reconstruction involves inferring both object geometry and spatial layout. Existing methods typically reconstruct objects independently or rely on implicit scene context, failing to exploit the repeated instances commonly present in realworld scenes. We propose FurnSet, a framework that explicitly identifies and leverages repeated object instances to improve reconstruction. Our method introduces per-object CLS tokens and a set-aware self-attention mechanism that groups identical instances and aggregates complementary observations across them, enabling joint reconstruction. We further combine scene-level and object-level conditioning to guide object reconstruction, followed by layout optimization using object point clouds with 3D and 2D projection losses for scene alignment. Experiments on 3D-Future and 3D-Front demonstrate improved scene reconstruction quality, highlighting the effectiveness of exploiting repetition for robust 3D scene reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。