arXiv:2410.11989cs.RO2024-10中稿 · ICRA被引 56

让机器人在动态环境中长期执行语言指令任务

Dynamic Open-Vocabulary 3D Scene Graphs for Long-term Language-Guided Mobile Manipulation

  • 用视觉语言模型识别物体并构建3D场景图
  • 局部更新场景图,无需重做整个重建
  • 适合需要持续适应变化环境的机器人任务

让移动机器人在频繁变化的真实环境中执行长期任务是一项重大挑战,尤其当环境因人机交互或机器人自身动作而不断改变时。传统方法通常假设场景静态,难以应对真实世界的动态性。为此,我们提出 DovSG 框架,结合动态开放词汇3D场景图与语言引导任务规划模块,实现长期任务执行。该框架输入RGB-D序列,利用视觉语言模型(VLMs)进行物体检测,获取高层次语义特征;基于分割后的物体生成结构化3D场景图以表达低层空间关系。此外,设计了一种高效的局部更新机制,使机器人在交互过程中可动态调整部分场景图,无需全场景重建。该机制在动态环境中尤为关键,能持续适应环境变化,有效支持长期任务。我们在不同人工修改程度的真实环境进行了验证,结果表明系统在长期任务中表现出色且性能优越。

原文摘要 · Abstract (English)

Enabling mobile robots to perform long-term tasks in dynamic real-world environments is a formidable challenge, especially when the environment changes frequently due to human-robot interactions or the robot's own actions. Traditional methods typically assume static scenes, which limits their applicability in the continuously changing real world. To overcome these limitations, we present DovSG, a novel mobile manipulation framework that leverages dynamic open-vocabulary 3D scene graphs and a language-guided task planning module for long-term task execution. DovSG takes RGB-D sequences as input and utilizes vision-language models (VLMs) for object detection to obtain high-level object semantic features. Based on the segmented objects, a structured 3D scene graph is generated for low-level spatial relationships. Furthermore, an efficient mechanism for locally updating the scene graph, allows the robot to adjust parts of the graph dynamically during interactions without the need for full scene reconstruction. This mechanism is particularly valuable in dynamic environments, enabling the robot to continually adapt to scene changes and effectively support the execution of long-term tasks. We validated our system in real-world environments with varying degrees of manual modifications, demonstrating its effectiveness and superior performance in long-term tasks. Our project page is available at: https://bjhyzj.github.io/dovsg-web.

机器人场景图动态环境语言引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。