arXiv:2411.15027cs.ROcs.AI2024-11被引 2

用大模型动态构建场景图,让机器人在变化环境中更智能地感知与协作。

Time is on my sight: scene graph filtering for dynamic environment perception in an LLM-driven robot

  • 通过RGB-D数据与粒子滤波生成实时更新的语义场景图
  • 大模型解析自然语言指令并规划导航、抓取等具体任务
  • 适合需要灵活交互的智能家居、医疗辅助机器人场景

机器人在工作场所、医院和家庭等动态环境中应用日益广泛,要求其感知能力能快速适应人类引起的环境变化。本文提出一种基于大语言模型(LLM)的机器人控制架构,解决人机交互中的核心挑战,重点在于动态创建并持续更新机器人状态表征。该架构利用LLM融合自然语言指令、机器人技能表示及实时动态语义地图,实现复杂动态环境下的灵活自适应行为。传统系统依赖静态预编程指令,难以应对实时变化。本架构通过感知模块基于RGB-D传感器数据生成并持续更新语义场景图,并采用粒子滤波确保动态环境中的精准物体定位。规划模块则利用最新语义地图将高层任务分解为子任务,关联导航、物体操作(如 PICK and PLACE)和移动(如 GOTO)等机器人技能。结合实时感知、状态追踪与大模型驱动的通信与任务规划,显著提升动态环境中的适应性、任务效率与人机协作能力。

原文摘要 · Abstract (English)

Robots are increasingly being used in dynamic environments like workplaces, hospitals, and homes. As a result, interactions with robots must be simple and intuitive, with robots perception adapting efficiently to human-induced changes. This paper presents a robot control architecture that addresses key challenges in human-robot interaction, with a particular focus on the dynamic creation and continuous update of the robot state representation. The architecture uses Large Language Models to integrate diverse information sources, including natural language commands, robotic skills representation, real-time dynamic semantic mapping of the perceived scene. This enables flexible and adaptive robotic behavior in complex, dynamic environments. Traditional robotic systems often rely on static, pre-programmed instructions and settings, limiting their adaptability to dynamic environments and real-time collaboration. In contrast, this architecture uses LLMs to interpret complex, high-level instructions and generate actionable plans that enhance human-robot collaboration. At its core, the system Perception Module generates and continuously updates a semantic scene graph using RGB-D sensor data, providing a detailed and structured representation of the environment. A particle filter is employed to ensure accurate object localization in dynamic, real-world settings. The Planner Module leverages this up-to-date semantic map to break down high-level tasks into sub-tasks and link them to robotic skills such as navigation, object manipulation (e.g., PICK and PLACE), and movement (e.g., GOTO). By combining real-time perception, state tracking, and LLM-driven communication and task planning, the architecture enhances adaptability, task efficiency, and human-robot collaboration in dynamic environments.

人机协作场景图大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。