综述大模型作为自主智能体和工具使用者的最新进展。
From Language to Action: A Review of Large Language Models as Autonomous Agents and Tool Users
- 系统梳理大模型在自主决策与工具调用中的架构设计
- 揭示大模型推理、规划与自我改进能力的关键机制
- 适合关注AI agent发展的研究者与工程师阅读
实现人类级人工智能的追求推动了自主智能体和大语言模型(LLMs)的发展。如今,LLMs被广泛用作决策智能体,因其具备理解指令、处理序列任务及通过反馈自适应的能力。本文综述2023至2025年间发表于A*、A类会议及Q1期刊的文献,围绕七个研究问题展开。分析了LLM智能体的架构设计原则,区分单智能体与多智能体系统,并探讨外部工具集成策略。同时考察了推理、规划、记忆等认知机制,以及提示方法与微调对智能体性能的影响。评估了现有基准与评测协议,并分析了68个公开数据集在不同任务中对基于LLM智能体的性能表现。研究发现大模型具备可验证推理能力、自我提升潜力与个性化定制空间。最后提出十项未来研究方向以弥补当前差距。
原文摘要 · Abstract (English)
The pursuit of human-level artificial intelligence (AI) has significantly advanced the development of autonomous agents and Large Language Models (LLMs). LLMs are now widely utilized as decision-making agents for their ability to interpret instructions, manage sequential tasks, and adapt through feedback. This review examines recent developments in employing LLMs as autonomous agents and tool users and comprises seven research questions. We only used the papers published between 2023 and 2025 in conferences of the A* and A rank and Q1 journals. A structured analysis of the LLM agents' architectural design principles, dividing their applications into single-agent and multi-agent systems, and strategies for integrating external tools is presented. In addition, the cognitive mechanisms of LLM, including reasoning, planning, and memory, and the impact of prompting methods and fine-tuning procedures on agent performance are also investigated. Furthermore, we evaluated current benchmarks and assessment protocols and have provided an analysis of 68 publicly available datasets to assess the performance of LLM-based agents in various tasks. In conducting this review, we have identified critical findings on verifiable reasoning of LLMs, the capacity for self-improvement, and the personalization of LLM-based agents. Finally, we have discussed ten future research directions to overcome these gaps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。