arXiv:2506.11102cs.CLcs.AI2025-06综述被引 16

厘清AI代理与大模型聊天机器人的评测差异,提供系统性框架。

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey

  • 从环境、反馈、感知等五方面区分AI代理与传统大模型。
  • 按外部驱动力和内部能力分类现有评测基准,构建参考表格。
  • 适合关注智能体评估的科研人员快速定位合适评测工具。

大型语言模型(如GPT、Gemini、DeepSeek)的兴起显著推动了自然语言处理发展,催生出能执行多样化语言任务的复杂聊天机器人。从传统大模型聊天机器人向更高级的AI代理演进,标志着关键转折。然而,现有评估框架常混淆二者界限,导致研究者在选择评测基准时困惑。本文基于演化视角,系统分析当前评估方法,提出一个清晰区分AI代理与大模型聊天机器人的五维框架:复杂环境、多源指令、动态反馈、多模态感知与高级能力。进一步,根据外部环境驱动力与衍生内部能力,对现有评测基准进行分类,并为每类明确相关评估属性,形成实用参考表。最后,从环境、代理、评估者与度量四个关键维度,总结趋势并展望未来评估方法。研究成果为研究者选择与应用评测基准提供可操作指导,助力该快速演进领域持续发展。

原文摘要 · Abstract (English)

The advent of large language models (LLMs), such as GPT, Gemini, and DeepSeek, has significantly advanced natural language processing, giving rise to sophisticated chatbots capable of diverse language-related tasks. The transition from these traditional LLM chatbots to more advanced AI agents represents a pivotal evolutionary step. However, existing evaluation frameworks often blur the distinctions between LLM chatbots and AI agents, leading to confusion among researchers selecting appropriate benchmarks. To bridge this gap, this paper introduces a systematic analysis of current evaluation approaches, grounded in an evolutionary perspective. We provide a detailed analytical framework that clearly differentiates AI agents from LLM chatbots along five key aspects: complex environment, multi-source instructor, dynamic feedback, multi-modal perception, and advanced capability. Further, we categorize existing evaluation benchmarks based on external environments driving forces, and resulting advanced internal capabilities. For each category, we delineate relevant evaluation attributes, presented comprehensively in practical reference tables. Finally, we synthesize current trends and outline future evaluation methodologies through four critical lenses: environment, agent, evaluator, and metrics. Our findings offer actionable guidance for researchers, facilitating the informed selection and application of benchmarks in AI agent evaluation, thus fostering continued advancement in this rapidly evolving research domain.

AI代理评测框架大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。