arXiv:2503.16416cs.AIcs.CL2025-03综述被引 216

首份LLM智能体评估综述,梳理评测方法与关键挑战

Survey on Evaluation of LLM-based Agents

  • 从五大维度系统分析智能体评估方法
  • 指出评测正向真实、动态、持续更新方向演进
  • 适合研究智能体评测与系统开发的学者参考

基于大语言模型的智能体代表了人工智能的新范式,使自主系统能够在动态环境中规划、推理并使用工具。本文首次全面综述了这类日益强大的智能体的评估方法。我们从五个视角进行分析:(1) 智能体工作流所需的核心大模型能力,如规划与工具使用;(2) 面向特定应用的基准测试,如网页操作和软件工程智能体;(3) 通用智能体的评估;(4) 智能体基准测试的核心维度分析;(5) 为开发者提供的评估框架与工具。分析揭示当前趋势:评测正转向更真实、更具挑战性且持续更新的基准。同时识别出未来研究亟需填补的关键空白,包括成本效益、安全性与鲁棒性的评估,以及细粒度、可扩展的评估方法。

原文摘要 · Abstract (English)

LLM-based agents represent a paradigm shift in AI, enabling autonomous systems to plan, reason, and use tools while interacting with dynamic environments. This paper provides the first comprehensive survey of evaluation methods for these increasingly capable agents. We analyze the field of agent evaluation across five perspectives: (1) Core LLM capabilities needed for agentic workflows, like planning, and tool use; (2) Application-specific benchmarks such as web and SWE agents; (3) Evaluation of generalist agents; (4) Analysis of agent benchmarks' core dimensions; and (5) Evaluation frameworks and tools for agent developers. Our analysis reveals current trends, including a shift toward more realistic, challenging evaluations with continuously updated benchmarks. We also identify critical gaps that future research must address, particularly in assessing cost-efficiency, safety, and robustness, and in developing fine-grained, scalable evaluation methods.

智能体评估大模型综述基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。