arXiv:2510.23357cs.RO2025-10综述被引 20

用大模型提升服务机器人任务规划能力,让机器人更智能地完成复杂家务。

Large language model-based task planning for service robots: A review

  • 将大模型作为机器人的认知核心,实现自主决策与规划。
  • 支持文本、视觉、语音等多模态输入,提升任务理解能力。
  • 适合研究智能机器人与AI融合的学者和工程师参考。

随着大语言模型(LLMs)和机器人技术的快速发展,服务机器人正日益融入日常生活,在复杂环境中提供多样化服务。为实现智能高效的服务交付,机器人需具备稳健准确的任务规划能力。本文全面综述了大模型在服务机器人中的集成应用,重点探讨其在增强机器人任务规划方面的作用。首先回顾了大模型的发展及基础技术,包括预训练、微调、检索增强生成(RAG)和提示工程。随后,分析大模型作为服务机器人认知核心(“大脑”)如何提升自主性与决策能力。进一步探讨了大模型在多种输入模态(文本、视觉、音频及多模态)下的任务规划进展。最后,总结当前研究的关键挑战与局限,并提出未来在复杂非结构化家庭环境中的发展方向。本综述旨在为人工智能与机器人领域的研究人员和实践者提供重要参考。

原文摘要 · Abstract (English)

With the rapid advancement of large language models (LLMs) and robotics, service robots are increasingly becoming an integral part of daily life, offering a wide range of services in complex environments. To deliver these services intelligently and efficiently, robust and accurate task planning capabilities are essential. This paper presents a comprehensive overview of the integration of LLMs into service robotics, with a particular focus on their role in enhancing robotic task planning. First, the development and foundational techniques of LLMs, including pre-training, fine-tuning, retrieval-augmented generation (RAG), and prompt engineering, are reviewed. We then explore the application of LLMs as the cognitive core-`brain'-of service robots, discussing how LLMs contribute to improved autonomy and decision-making. Furthermore, recent advancements in LLM-driven task planning across various input modalities are analyzed, including text, visual, audio, and multimodal inputs. Finally, we summarize key challenges and limitations in current research and propose future directions to advance the task planning capabilities of service robots in complex, unstructured domestic environments. This review aims to serve as a valuable reference for researchers and practitioners in the fields of artificial intelligence and robotics.

大模型机器人任务规划多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。