用大模型实现能听懂人话、规划任务并反馈的机器人,成功率提升28%。
PFEA: An LLM-based High-Level Natural Language Planning and Feedback Embodied Agent for Human-Centered AI
- 分三模块:语音交互、视觉语言规划、动作执行
- 在仿真与真实环境任务成功率高出28%
- 适合做智能客服、家庭助手机器人
大型语言模型(LLMs)的快速发展标志着人工智能的新突破,开启了以人为本的人工智能(HAI)时代。HAI致力于更好服务人类福祉与需求,对机器人智能提出更高要求,尤其体现在自然语言交互、复杂任务规划与执行方面。基于LLM的智能体为实现HAI开辟了新路径。然而,现有基于LLM的具身智能体往往缺乏在线规划与执行复杂自然语言指令的能力。本文探索了基于视觉-语言模型(VLMs)的机器人操作智能体在物理世界中的实现。我们提出一种新型具身智能体框架,包含人机语音交互模块、视觉-语言智能体模块和动作执行模块。视觉-语言智能体自身包含基于视觉的任务规划器、自然语言指令转换器和任务执行反馈评估器。实验结果表明,相较于仅依赖LLM+CLIP的方法,本智能体在仿真与真实环境中平均任务成功率提高28%,显著提升了高层次自然语言指令任务的执行成功率。
原文摘要 · Abstract (English)
The rapid advancement of Large Language Models (LLMs) has marked a significant breakthrough in Artificial Intelligence (AI), ushering in a new era of Human-centered Artificial Intelligence (HAI). HAI aims to better serve human welfare and needs, thereby placing higher demands on the intelligence level of robots, particularly in aspects such as natural language interaction, complex task planning, and execution. Intelligent agents powered by LLMs have opened up new pathways for realizing HAI. However, existing LLM-based embodied agents often lack the ability to plan and execute complex natural language control tasks online. This paper explores the implementation of intelligent robotic manipulating agents based on Vision-Language Models (VLMs) in the physical world. We propose a novel embodied agent framework for robots, which comprises a human-robot voice interaction module, a vision-language agent module and an action execution module. The vision-language agent itself includes a vision-based task planner, a natural language instruction converter, and a task performance feedback evaluator. Experimental results demonstrate that our agent achieves a 28\% higher average task success rate in both simulated and real environments compared to approaches relying solely on LLM+CLIP, significantly improving the execution success rate of high-level natural language instruction tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。