arXiv:2603.15483cs.AI2026-03中稿 · ICLR被引 2

新框架自动评估智能体,兼顾用户角色与错误诊断。

Talk, Evaluate, Diagnose: User-aware Agent Evaluation with Automated Error Analysis

  • 用通用用户角色模板模拟专家与非专家交互。
  • 引入新指标,提升对对话效率与进展的评估精度。
  • 自动化分析错误模式,助力智能体性能提升8%-10%。

智能体应用正被广泛用于自动化各类任务,但其运行领域多样,难以建立可扩展的统一评估框架。现有方法各自定义任务成功标准(如数据库查询、正则匹配),增加开发复杂性,且未系统考虑用户角色与专业度,导致评估信息不完整。本文提出TED框架(Talk, Evaluate, Diagnose):(1) Talk——使用可复用的专家与非专家用户角色模板进行交互;(2) Evaluate——将子目标(如工具签名、响应)转化为自然语言评分说明,由大模型作为评判者自动评估,并提出融合对话轮次效率与中间进展的新指标;(3) Diagnose——开发自动化错误分析工具,识别评判者与智能体间的不一致,揭示常见错误并提供可操作反馈。实验表明,该框架在不同模型与用户专业度下揭示了新的性能洞察,结合修复建议后,智能体性能在新指标上最高提升8%-10%。

原文摘要 · Abstract (English)

Agent applications are increasingly adopted to automate workflows across diverse tasks. However, due to the heterogeneous domains they operate in, it is challenging to create a scalable evaluation framework. Prior works each employ their own methods to determine task success, such as database lookups, regex match, etc., adding complexity to the development of a unified agent evaluation approach. Moreover, they do not systematically account for the user's role nor expertise in the interaction, providing incomplete insights into the agent's performance. We argue that effective agent evaluation goes beyond correctness alone, incorporating conversation quality, efficiency and systematic diagnosis of agent errors. To address this, we introduce the TED framework (Talk, Evaluate, Diagnose). (1) Talk: We leverage reusable, generic expert and non-expert user persona templates for user-agent interaction. (2) Evaluate: We adapt existing datasets by representing subgoals-such as tool signatures, and responses-as natural language grading notes, evaluated automatically with LLM-as-a-judge. We propose new metrics that capture both turn efficiency and intermediate progress of the agent complementing the user-aware setup. (3) Diagnose: We introduce an automated error analysis tool that analyzes the inconsistencies of the judge and agents, uncovering common errors, and providing actionable feedback for agent improvement. We show that our TED framework reveals new insights regarding agent performance across models and user expertise levels. We also demonstrate potential gains in agent performance with peaks of 8-10% on our proposed metrics after incorporating the identified error remedies into the agent's design.

智能体评估用户角色错误诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。