arXiv:2504.19678cs.AIcs.LG2025-04综述被引 213

系统梳理大模型到自主智能体的评估与协作体系,构建统一框架。

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review

  • 整合2019-2025年60个评测基准,覆盖多领域任务
  • 提出涵盖数学、编程、科学推理等的分类体系
  • 适合关注智能体评估与协同研究的研究者

大型语言模型与自主智能体发展迅速,催生了多样化的评估基准、框架与协作协议。为应对标准化评估与集成需求,本文系统整合这些分散工作,构建统一框架。尽管如此,当前领域仍碎片化,缺乏统一分类体系与全面综述。为此,我们对比分析了2019至2025年间开发的多个评测基准,涵盖通用与学术知识推理、数学问题求解、代码生成与软件工程、事实对齐与检索、领域特定任务、多模态与具身任务、任务编排及交互评估。同时,回顾2023至2025年引入的智能体框架,整合大模型与模块化工具包以实现自主决策与多步推理。进一步展示智能体在材料科学、生物医学研究、学术创意生成、软件工程、合成数据生成、化学推理、数学求解、地理信息系统、多媒体、医疗健康与金融等领域的实际应用。还调研了关键的智能体间协作协议,包括代理通信协议(ACP)、模型上下文协议(MCP)与代理-代理协议(A2A)。最后,讨论未来研究方向,聚焦高级推理策略、多智能体系统的失败模式、自动化科学发现、基于强化学习的动态工具集成、集成搜索能力以及智能体协议中的安全漏洞。

原文摘要 · Abstract (English)

Large language models and autonomous AI agents have evolved rapidly, resulting in a diverse array of evaluation benchmarks, frameworks, and collaboration protocols. Driven by the growing need for standardized evaluation and integration, we systematically consolidate these fragmented efforts into a unified framework. However, the landscape remains fragmented and lacks a unified taxonomy or comprehensive survey. Therefore, we present a side-by-side comparison of benchmarks developed between 2019 and 2025 that evaluate these models and agents across multiple domains. In addition, we propose a taxonomy of approximately 60 benchmarks that cover general and academic knowledge reasoning, mathematical problem-solving, code generation and software engineering, factual grounding and retrieval, domain-specific evaluations, multimodal and embodied tasks, task orchestration, and interactive assessments. Furthermore, we review AI-agent frameworks introduced between 2023 and 2025 that integrate large language models with modular toolkits to enable autonomous decision-making and multi-step reasoning. Moreover, we present real-world applications of autonomous AI agents in materials science, biomedical research, academic ideation, software engineering, synthetic data generation, chemical reasoning, mathematical problem-solving, geographic information systems, multimedia, healthcare, and finance. We then survey key agent-to-agent collaboration protocols, namely the Agent Communication Protocol (ACP), the Model Context Protocol (MCP), and the Agent-to-Agent Protocol (A2A). Finally, we discuss recommendations for future research, focusing on advanced reasoning strategies, failure modes in multi-agent LLM systems, automated scientific discovery, dynamic tool integration via reinforcement learning, integrated search capabilities, and security vulnerabilities in agent protocols.

智能体大模型评估基准协作协议

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。