多机器人通过3D场景图理解自然语言指令并协同执行复杂任务。
Language-Grounded Hierarchical Planning and Execution with Multi-Robot 3D Scene Graphs
- 用3D场景图融合多机器人感知数据,实现共享环境认知。
- 结合LLM与场景图,将自然语言指令转化为可执行的规划目标。
- 系统在大型户外真实环境中验证,支持实时重定位与任务执行。
本文提出一种多机器人系统,整合建图、定位与任务运动规划(TAMP),基于3D场景图执行自然语言表达的复杂指令。系统构建共享3D场景图,包含开放集物体地图,用于多机器人3D场景图融合。该表示支持基于物体地图的实时、视角不变重定位,以及基于3D场景图的规划,使机器人团队能够推理环境并执行复杂任务。此外,我们引入一种规划方法,利用大语言模型(LLM)结合共享3D场景图和机器人能力上下文,将操作员意图转换为规划领域定义语言(PDDL)目标。我们在大规模室外真实环境中评估了系统的性能。补充视频见 https://youtu.be/8xbGGOLfLAY。
原文摘要 · Abstract (English)
In this paper, we introduce a multi-robot system that integrates mapping, localization, and task and motion planning (TAMP) enabled by 3D scene graphs to execute complex instructions expressed in natural language. Our system builds a shared 3D scene graph incorporating an open-set object-based map, which is leveraged for multi-robot 3D scene graph fusion. This representation supports real-time, view-invariant relocalization (via the object-based map) and planning (via the 3D scene graph), allowing a team of robots to reason about their surroundings and execute complex tasks. Additionally, we introduce a planning approach that translates operator intent into Planning Domain Definition Language (PDDL) goals using a Large Language Model (LLM) by leveraging context from the shared 3D scene graph and robot capabilities. We provide an experimental assessment of the performance of our system on real-world tasks in large-scale, outdoor environments. A supplementary video is available at https://youtu.be/8xbGGOLfLAY.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。