arXiv:2504.01252cs.ROcs.AI2025-04被引 1

用大模型实现机器人在人机互动中合理规划行动,成功率超90%。

Plan-and-Act using Large Language Models for Interactive Agreement

  • 给大模型提供当前机器人动作文本,实现被动与主动交互的切换。
  • 引入第二阶段提问判断何时调用大模型,确保行动时机恰当。
  • 该方法可快速适配多种人机交互场景,适合需要智能协作的机器人应用。

近期大型语言模型(LLMs)已具备规划机器人动作的能力。本文探讨如何利用大模型处理涉及情境化人机交互(HRI)的任务。关键挑战在于平衡‘尊重人类当前活动’与‘优先完成机器人任务’之间的关系,以及确定使用大模型生成动作计划的时机。为此,本文提出一种必要的‘规划-执行’技能设计:通过向大模型提供当前机器人的动作文本,使其能准确判断是否介入;同时引入第二阶段提问,以决定下一次调用大模型的最佳时机。该设计应用于‘参与’(Engage)技能,并在四个不同交互场景中测试,结果显示,采用该方法后,大模型可有效扩展至多种场景,测试成功率高达90%。

原文摘要 · Abstract (English)

Recent large language models (LLMs) are capable of planning robot actions. In this paper, we explore how LLMs can be used for planning actions with tasks involving situational human-robot interaction (HRI). A key problem of applying LLMs in situational HRI is balancing between "respecting the current human's activity" and "prioritizing the robot's task," as well as understanding the timing of when to use the LLM to generate an action plan. In this paper, we propose a necessary plan-and-act skill design to solve the above problems. We show that a critical factor for enabling a robot to switch between passive / active interaction behavior is to provide the LLM with an action text about the current robot's action. We also show that a second-stage question to the LLM (about the next timing to call the LLM) is necessary for planning actions at an appropriate timing. The skill design is applied to an Engage skill and is tested on four distinct interaction scenarios. We show that by using the skill design, LLMs can be leveraged to easily scale to different HRI scenarios with a reasonable success rate reaching 90% on the test scenarios.

人机交互大模型机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。