arXiv:2409.10350cs.ROcs.AI2024-09ICRA被引 11

仅用点云生成3D开放词汇场景图,无需相机图像或位姿。

Point2Graph: An End-to-end Point Cloud-based 3D Open-Vocabulary Scene Graph for Robot Navigation

  • 基于点云的分层框架,融合几何与学习方法分割房间和物体。
  • 在真实场景数据集上超越当前最先进算法的物体与房间分类性能。
  • 适合无摄像头或位姿信息的机器人导航场景,如未知环境探索。

现有开放词汇场景图生成算法高度依赖3D场景点云和带位姿的RGB-D图像,因此在缺乏RGB-D图像或相机位姿的场景中应用受限。为解决此问题,我们提出Point2Graph,一种全新的端到端点云驱动的3D开放词汇场景图生成框架,不再需要带位姿的RGB-D图像序列。该分层框架包含房间与物体检测/分割及开放词汇分类。对于房间层,我们结合基于几何的边界检测与基于学习的区域检测,实现房间分割,并构建了‘Snap-Lookup’框架用于开放词汇房间分类。此外,我们设计了端到端的物体层管道,仅基于3D点云数据完成物体检测与分类。评估结果表明,该框架在广泛使用的实际场景数据集上,优于当前最先进的开放词汇物体与房间分割及分类算法。

原文摘要 · Abstract (English)

Current open-vocabulary scene graph generation algorithms highly rely on both 3D scene point cloud data and posed RGB-D images and thus have limited applications in scenarios where RGB-D images or camera poses are not readily available. To solve this problem, we propose Point2Graph, a novel end-to-end point cloud-based 3D open-vocabulary scene graph generation framework in which the requirement of posed RGB-D image series is eliminated. This hierarchical framework contains room and object detection/segmentation and open-vocabulary classification. For the room layer, we leverage the advantage of merging the geometry-based border detection algorithm with the learning-based region detection to segment rooms and create a "Snap-Lookup" framework for open-vocabulary room classification. In addition, we create an end-to-end pipeline for the object layer to detect and classify 3D objects based solely on 3D point cloud data. Our evaluation results show that our framework can outperform the current state-of-the-art (SOTA) open-vocabulary object and room segmentation and classification algorithm on widely used real-scene datasets.

3D场景理解点云处理机器人导航开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。