让机器人理解任务状态,生成可执行的协作计划。
STEP: State-Aware Task Estimation and Planning with Multi-Modal LLMs for Human-Robot Collaboration

- 用多模态大模型显式预测系统状态与动作后的状态变化
- 在仿真中使动作可执行性提升32.8%,最终状态误差降低14.8%
- 适合需要精准状态感知的工业人机协作场景
工业场景中有效的人机协作要求机器人理解人类意图并辅助任务规划以减轻负担。现有方法利用多模态大语言模型(MM-LLMs)在数据稀缺情况下通过上下文学习解析用户行为,并生成自然语言描述的长时序动作计划。然而,MM-LLMs缺乏对系统状态的理解,无法跟踪状态转移,常导致幻觉动作偏离目标。且自然语言生成的计划层级过高,执行时存在歧义。为此,我们提出状态感知任务估计与规划器(STEP),通过提示MM-LLM显式估计系统状态,并预测动作引发的状态转移。通过同时预测未来状态与动作,STEP确保任务收敛性,并提供执行所需额外辅助参数。我们在模拟机器人装配任务环境中评估,结果表明该方法在动作可执行性上优于当前最优方法32.8%,最终状态误差降低14.8%。
原文摘要 · Abstract (English)
Effective human-robot collaboration in industrial settings requires robots to understand human intentions and assist with task planning, reducing workload. Recent works have explored the use of Multi-modal Large Language Models (MM-LLMs) for task planning in such data-scarce scenarios, leveraging in-context learning to interpret user actions and generate long-horizon action plans in natural language. However, MM-LLMs inherently lack an understanding of system states and do not track state transitions, often leading to hallucinated actions that deviate from the intended goal. Additionally, generating action plans in natural language tends to limit the generated plans to a high level, introducing ambiguity in action execution. To address these limitations, we propose the State-aware Task Estimator and Planner (STEP), which prompts a MM-LLM to explicitly estimate the state of the system and predict the state transitions resulting from executed actions. By forecasting future states alongside actions, STEP ensures task-convergent planning while also providing additional assistance parameters necessary for executing the predicted actions. We evaluate STEP in a simulated environment using a robot assembly task. Our approach outperforms the state-of-the-art by 32.8% in action executability and 14.8% in final-state error.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。