整合大模型、知识库与推理能力,构建下一代具身智能代理
Towards general embodied intelligence: integrating large language models, knowledge bases, and reasoning capabilities to build the next generation of AI agents

- 将大语言模型与知识库、逻辑推理结合,推动具身智能发展
- 提出四维协同框架,涵盖感知、推理与行动的统一建模
- 面向复杂动态环境,适合研究通用人工智能与智能体系统
大语言模型(LLMs)、结构化知识库(KBs)与推理能力(RA)的融合,为通用具身智能(GEI)提供了重要路径。本文回顾以大模型为核心的智能系统演进,强调其与知识表征、逻辑推理及物理具身的结合。分析了大模型架构、预训练方法与推理机制,及其与外部知识源和结构化推理框架的交互。同时探讨了智能体在物理环境中学习与行动的具身智能范式。为整合上述维度,提出一个概念性框架,展示大模型、知识库、推理能力与具身性的协同作用,作为感知、推理与行动的指导模型,而非具体工程实现。为推进通用具身智能,识别五大挑战:高效大模型部署、闭环知识整合、符号-神经混合推理、感知-动作对齐、持续学习。本综述为开发适应复杂动态环境的多模态自适应智能体提供了全面路线图。
原文摘要 · Abstract (English)
The convergence of large language models (LLMs), structured knowledge bases (KBs), and reasoning ability (RA) presents a promising trajectory toward general embodied intelligence (GEI). This paper reviews the evolution of LLM-centered intelligent systems, emphasising their integration with knowledge representation, logical reasoning, and physical embodiment. We analyse LLM architectures, pre-training methods, and inference mechanisms, along with their interaction with external knowledge sources and structured reasoning frameworks. Furthermore, we examine embodied intelligence (EI) paradigms wherein agents learn and act in physical environments. To synthesise these dimensions, we present a conceptual framework that illustrates the synergy among LLMs, KBs, RA, and embodiment, serving as a guiding model for perception, reasoning, and action rather than an implemented engineering architecture. To advance toward GEI, we identify five key challenges: efficient LLM deployment, closed-loop knowledge integration, hybrid symbolic-neural reasoning, perception-action grounding, and continual learning. This survey provides a comprehensive roadmap for developing adaptive, multimodal agents capable of operating in complex, dynamic settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。