用户用手移动物体,系统自动识别并重建每个物体的3D模型。
PickScan: Object discovery and reconstruction from handheld interactions
- 通过物体位移检测交互,不依赖物体类别先验。
- 99%召回率下精度78.3%,平均切比雪夫距离0.90厘米。
- 适合需要手动操作的复杂场景重建,如机器人抓取与增强现实。
构建场景中每个物体独立的3D表示是机器人和增强现实中的理想能力。然而,现有方法大多依赖强外观先验,仅适用于训练过的物体类别,或无法支持物体操作,难以完整扫描并引导复杂场景下的物体发现。本文提出一种新型交互引导、类无关的方法,基于物体位移实现用户手持RGB-D相机移动场景、拿起物体,最终为每个被拿起的物体输出一个3D模型。核心贡献在于一种新交互检测与操纵物体掩码提取方法。在自建数据集上,本方法在100%召回率下达到78.3%精度,并以0.90厘米的平均切比雪夫距离完成重建。相比唯一可比的交互式、类无关基线Co-Fusion,切比雪夫距离降低73%,误报减少99%。
原文摘要 · Abstract (English)
Reconstructing compositional 3D representations of scenes, where each object is represented with its own 3D model, is a highly desirable capability in robotics and augmented reality. However, most existing methods rely heavily on strong appearance priors for object discovery, therefore only working on those classes of objects on which the method has been trained, or do not allow for object manipulation, which is necessary to scan objects fully and to guide object discovery in challenging scenarios. We address these limitations with a novel interaction-guided and class-agnostic method based on object displacements that allows a user to move around a scene with an RGB-D camera, hold up objects, and finally outputs one 3D model per held-up object. Our main contribution to this end is a novel approach to detecting user-object interactions and extracting the masks of manipulated objects. On a custom-captured dataset, our pipeline discovers manipulated objects with 78.3% precision at 100% recall and reconstructs them with a mean chamfer distance of 0.90cm. Compared to Co-Fusion, the only comparable interaction-based and class-agnostic baseline, this corresponds to a reduction in chamfer distance of 73% while detecting 99% fewer false positives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。