arXiv:2503.07909cs.CVcs.AI2025-03中稿 · IROS 2025被引 22

让机器人通过细粒度功能部件理解环境,实现精准交互。

FunGraph: Functionality Aware 3D Scene Graphs for Language-Prompted Scene Interaction

  • 在3D场景图中加入功能部件定位与使用方式信息
  • 实测功能部件分割精度达当前最优水平
  • 适合需要精准环境交互的机器人研究者

3D场景图作为环境的语义与层次化表示正日益受到重视。现有方法多停留在粗粒度的对象级层面。本文目标是构建一种可直接支持机器人交互的表示,识别功能交互部件的位置及其使用方式。为此,我们聚焦于以更细粒度检测并存储与可及性相关的部件。主要挑战在于缺乏超越实例检测的数据,以及机器人传感器难以捕捉详细物体特征。我们利用现有3D资源生成2D数据并训练检测器,进而增强标准3D场景图生成流程。实验表明,本方法在功能部件分割上达到与先进3D模型相当的性能,且其增强方案在任务驱动的可及性定位上显著优于现有方法。

原文摘要 · Abstract (English)

The concept of 3D scene graphs is increasingly recognized as a powerful semantic and hierarchical representation of the environment. Current approaches often address this at a coarse, object-level resolution. In contrast, our goal is to develop a representation that enables robots to directly interact with their environment by identifying both the location of functional interactive elements and how these can be used. To achieve this, we focus on detecting and storing objects at a finer resolution, focusing on affordance-relevant parts. The primary challenge lies in the scarcity of data that extends beyond instance-level detection and the inherent difficulty of capturing detailed object features using robotic sensors. We leverage currently available 3D resources to generate 2D data and train a detector, which is then used to augment the standard 3D scene graph generation pipeline. Through our experiments, we demonstrate that our approach achieves functional element segmentation comparable to state-of-the-art 3D models and that our augmentation enables task-driven affordance grounding with higher accuracy than the current solutions. See our project page at https://fungraph.github.io.

3D场景图机器人交互功能部件

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。