arXiv:2503.22458cs.CLcs.AI2025-03综述被引 68

系统梳理大模型对话智能体的评估方法,构建双维度分类框架。

Evaluating LLM-based Agents for Multi-Turn Conversations: A Survey

  • 构建‘评估什么’与‘如何评估’双维度分类体系
  • 涵盖任务完成、响应质量、记忆保持等核心维度
  • 适合研究对话系统评估或开发智能代理的研究者

本综述系统考察了大语言模型(LLM)驱动的智能体在多轮对话场景中的评估方法。基于类PRISMA框架,我们系统梳理了近250篇学术文献,覆盖多个出版渠道,为分析奠定基础。研究提出两个相互关联的分类体系:其一定义评估对象,包括任务完成度、响应质量、用户体验、记忆与上下文保持、规划与工具集成等关键组件,确保评估全面且有意义;其二聚焦评估方法,将手段分为标注式评估、自动化指标、人机结合的混合策略,以及利用大模型自评的自判断方法。该框架不仅包含传统语言理解指标(如BLEU、ROUGE),还纳入反映多轮对话动态交互特性的先进评估技术。

原文摘要 · Abstract (English)

This survey examines evaluation methods for large language model (LLM)-based agents in multi-turn conversational settings. Using a PRISMA-inspired framework, we systematically reviewed nearly 250 scholarly sources, capturing the state of the art from various venues of publication, and establishing a solid foundation for our analysis. Our study offers a structured approach by developing two interrelated taxonomy systems: one that defines \emph{what to evaluate} and another that explains \emph{how to evaluate}. The first taxonomy identifies key components of LLM-based agents for multi-turn conversations and their evaluation dimensions, including task completion, response quality, user experience, memory and context retention, as well as planning and tool integration. These components ensure that the performance of conversational agents is assessed in a holistic and meaningful manner. The second taxonomy system focuses on the evaluation methodologies. It categorizes approaches into annotation-based evaluations, automated metrics, hybrid strategies that combine human assessments with quantitative measures, and self-judging methods utilizing LLMs. This framework not only captures traditional metrics derived from language understanding, such as BLEU and ROUGE scores, but also incorporates advanced techniques that reflect the dynamic, interactive nature of multi-turn dialogues.

大模型评估对话系统智能体综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。