用大模型让机器人更智能,能听懂指令、自己规划任务
Towards Embodied Agentic AI: Review and Classification of LLM- and VLM-Driven Robot Autonomy and Interaction
- 用大语言模型和视觉语言模型驱动机器人自主决策
- 支持自然语言理解、任务规划与多步操作执行
- 适合研究人机交互与智能机器人系统的开发者
基础模型,包括大语言模型(LLMs)和视觉-语言模型(VLMs),最近推动了机器人自主性和人机接口的新范式。同时,视觉-语言-动作模型(VLAs)或大行为模型(LBMs)正提升机器人的灵巧性与能力。本文综述了促进代理型应用与架构的研究工作,涵盖早期基于GPT风格的接口到更复杂的系统,其中AI代理作为协调者、规划者、感知执行者或通用接口。此类代理架构使机器人能够基于自然语言指令进行推理、调用API、规划任务序列或协助运维与诊断。除同行评审论文外,鉴于该领域发展迅速,本文还纳入社区项目、ROS包及工业框架,以反映新兴趋势。我们提出了一个模型集成方式的分类体系,并对当前文献中代理角色进行了对比分析。
原文摘要 · Abstract (English)
Foundation models, including large language models (LLMs) and vision-language models (VLMs), have recently enabled novel approaches to robot autonomy and human-robot interfaces. In parallel, vision-language-action models (VLAs) or large behavior models (LBMs) are increasing the dexterity and capabilities of robotic systems. This survey paper reviews works that advance agentic applications and architectures, including initial efforts with GPT-style interfaces and more complex systems where AI agents function as coordinators, planners, perception actors, or generalist interfaces. Such agentic architectures allow robots to reason over natural language instructions, invoke APIs, plan task sequences, or assist in operations and diagnostics. In addition to peer-reviewed research, due to the fast-evolving nature of the field, we highlight and include community-driven projects, ROS packages, and industrial frameworks that show emerging trends. We propose a taxonomy for classifying model integration approaches and present a comparative analysis of the role that agents play in different solutions in today's literature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。