梳理大模型智能体评估体系,帮人判断其真实能力与安全可靠性。
Evaluation and Benchmarking of LLM Agents: A Survey

- 按评估目标和流程构建双维度分类框架
- 指出企业部署中可靠性、合规性等关键挑战
- 适合研究者和工程师系统评估智能体落地可行性
大模型智能体的兴起拓展了人工智能应用边界,但其评估仍复杂且发展不足。本综述深入梳理智能体评估新兴领域,提出二维分类框架:一是评估目标(如智能体行为、能力、可靠性与安全性),二是评估流程(包括交互模式、数据集与基准、指标计算方法及工具)。除分类外,还强调企业级挑战,如基于角色的数据访问、可靠性保障、长周期动态交互与合规要求,这些常被现有研究忽视。同时指明未来方向:更全面、更真实、可扩展的评估体系。本文旨在厘清评估领域的碎片化现状,提供系统性评估框架,助力研究者与实践者为真实场景部署评估大模型智能体。
原文摘要 · Abstract (English)
The rise of LLM-based agents has opened new frontiers in AI applications, yet evaluating these agents remains a complex and underdeveloped area. This survey provides an in-depth overview of the emerging field of LLM agent evaluation, introducing a two-dimensional taxonomy that organizes existing work along (1) evaluation objectives -- what to evaluate, such as agent behavior, capabilities, reliability, and safety -- and (2) evaluation process -- how to evaluate, including interaction modes, datasets and benchmarks, metric computation methods, and tooling. In addition to taxonomy, we highlight enterprise-specific challenges, such as role-based access to data, the need for reliability guarantees, dynamic and long-horizon interactions, and compliance, which are often overlooked in current research. We also identify future research directions, including holistic, more realistic, and scalable evaluation. This work aims to bring clarity to the fragmented landscape of agent evaluation and provide a framework for systematic assessment, enabling researchers and practitioners to evaluate LLM agents for real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。