arXiv:2602.02456cs.RO2026-02被引 4

用分层3D场景图让机器人理解物体关系并完成任务推理。

Relationship-Aware Hierarchical 3D Scene Graph for Task Reasoning

  • 构建多层级3D场景图,融合视觉语言模型的语义特征。
  • 引入任务推理模块,结合LLM与VLM实现智能决策。
  • 在四足机器人上验证,支持复杂环境下的任务理解。

以结构化方式表征和理解3D环境对自主代理导航与环境推理至关重要。传统同步定位与地图构建(SLAM)方法虽可生成度量重建并扩展为度量-语义地图,但缺乏高层抽象与关系推理能力。为此,3D场景图成为捕捉层次结构与物体关系的有力工具。本文提出一种增强的分层3D场景图,整合跨多个抽象层级的开放词汇特征,并支持物体间关系推理。方法利用视觉语言模型(VLM)推断语义关系,特别引入任务推理模块,结合大语言模型(LLM)与VLM,解析场景图中的语义与关系信息,使代理能更智能地进行任务推理与环境交互。我们在多种环境与任务中部署该方法于四足机器人,验证其任务推理能力。

原文摘要 · Abstract (English)

Representing and understanding 3D environments in a structured manner is crucial for autonomous agents to navigate and reason about their surroundings. While traditional Simultaneous Localization and Mapping (SLAM) methods generate metric reconstructions and can be extended to metric-semantic mapping, they lack a higher level of abstraction and relational reasoning. To address this gap, 3D scene graphs have emerged as a powerful representation for capturing hierarchical structures and object relationships. In this work, we propose an enhanced hierarchical 3D scene graph that integrates open-vocabulary features across multiple abstraction levels and supports object-relational reasoning. Our approach leverages a Vision Language Model (VLM) to infer semantic relationships. Notably, we introduce a task reasoning module that combines Large Language Models (LLM) and a VLM to interpret the scene graph's semantic and relational information, enabling agents to reason about tasks and interact with their environment more intelligently. We validate our method by deploying it on a quadruped robot in multiple environments and tasks, highlighting its ability to reason about them.

3D场景图关系推理机器人多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。