大模型赋能具身智能,提升决策与学习能力
Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- 用大模型增强具身系统感知、规划与执行能力
- 整合世界模型,改进自主决策与学习效率
- 适合关注AGI与机器人智能的研究者阅读
具身人工智能致力于开发具备物理形态的智能系统,使其能在真实环境中感知、决策、行动并学习,是通往通用人工智能(AGI)的重要路径。尽管历经数十年探索,具身代理在开放动态环境中的通用任务仍难以达到人类水平。近年来大模型的突破显著推动了具身智能的发展,提升了感知、交互、规划与学习能力。本文全面综述大模型赋能的具身智能,聚焦自主决策与具身学习。我们分析分层与端到端决策范式,阐明大模型如何增强高层规划、底层执行与反馈机制;同时探讨其对视觉-语言-动作(VLA)模型的改进作用。针对具身学习,介绍主流方法,深入分析大模型在模仿学习与强化学习中的增强效果。首次将世界模型纳入具身智能综述,阐述其设计方法与在决策与学习中的关键作用。尽管已有显著进展,仍存在诸多挑战,文中最后讨论这些瓶颈及其可能的研究方向。
原文摘要 · Abstract (English)
Embodied AI aims to develop intelligent systems with physical forms capable of perceiving, decision-making, acting, and learning in real-world environments, providing a promising way to Artificial General Intelligence (AGI). Despite decades of explorations, it remains challenging for embodied agents to achieve human-level intelligence for general-purpose tasks in open dynamic environments. Recent breakthroughs in large models have revolutionized embodied AI by enhancing perception, interaction, planning and learning. In this article, we provide a comprehensive survey on large model empowered embodied AI, focusing on autonomous decision-making and embodied learning. We investigate both hierarchical and end-to-end decision-making paradigms, detailing how large models enhance high-level planning, low-level execution, and feedback for hierarchical decision-making, and how large models enhance Vision-Language-Action (VLA) models for end-to-end decision making. For embodied learning, we introduce mainstream learning methodologies, elaborating on how large models enhance imitation learning and reinforcement learning in-depth. For the first time, we integrate world models into the survey of embodied AI, presenting their design methods and critical roles in enhancing decision-making and learning. Though solid advances have been achieved, challenges still exist, which are discussed at the end of this survey, potentially as the further research directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。