用大模型提升机器人理解与决策能力,推动智能机器人发展。
Foundation Model Driven Robotics: A Comprehensive Review
- 利用大语言和视觉语言模型实现跨模态推理与高层规划
- 验证了在仿真到现实迁移中的泛化能力与系统集成效果
- 适合关注机器人智能化、多模态融合的科研与工程人员
大型语言模型(LLMs)和视觉-语言模型(VLMs)的快速发展为机器人技术带来了范式变革。这些模型具备强大的语义理解、高层推理和跨模态泛化能力,显著推动了感知、规划、控制及人机交互的进步。本文系统综述了近年来在仿真驱动设计、开放世界执行、仿真到现实迁移以及可适应机器人领域的进展,强调集成化、系统级策略,并评估其在真实环境中的可行性。文章讨论了过程化场景生成、策略泛化与多模态推理等关键趋势,同时指出局限性:具身能力有限、多模态数据不足、安全风险与计算约束。基于此,论文揭示了基础模型在实时性、对齐性、鲁棒性和可信度方面的挑战,并提出未来研究路线图,旨在通过更稳健、可解释、具身化的模型,弥合语义推理与物理智能之间的鸿沟。
原文摘要 · Abstract (English)
The rapid emergence of foundation models, particularly Large Language Models (LLMs) and Vision-Language Models (VLMs), has introduced a transformative paradigm in robotics. These models offer powerful capabilities in semantic understanding, high-level reasoning, and cross-modal generalization, enabling significant advances in perception, planning, control, and human-robot interaction. This critical review provides a structured synthesis of recent developments, categorizing applications across simulation-driven design, open-world execution, sim-to-real transfer, and adaptable robotics. Unlike existing surveys that emphasize isolated capabilities, this work highlights integrated, system-level strategies and evaluates their practical feasibility in real-world environments. Key enabling trends such as procedural scene generation, policy generalization, and multimodal reasoning are discussed alongside core bottlenecks, including limited embodiment, lack of multimodal data, safety risks, and computational constraints. Through this lens, this paper identifies both the architectural strengths and critical limitations of foundation model-based robotics, highlighting open challenges in real-time operation, grounding, resilience, and trust. The review concludes with a roadmap for future research aimed at bridging semantic reasoning and physical intelligence through more robust, interpretable, and embodied models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。