arXiv:2605.22608cs.CLcs.AI2026-05ACL被引 2

自动评估大模型智能体行为,三层次洞察其表现。

Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents

论文配图:Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents
图 1 · 摘自论文原文
  • 动态生成多粒度评估报告,覆盖系统、执行轨迹与节点层级。
  • 在四基准、七场景下实现与人工标注高度一致的错误识别。
  • 界面直观易用,适合开发者快速诊断智能体缺陷。

随着智能体具备自主制定策略、执行动作并交互环境的能力,对其行为的监督与评估面临严峻挑战。现有工具多局限于基础可观测性或依赖静态手工错误分类体系,难以适应新领域。为此,我们提出 Agentic CLEAR——一个自动、动态且易于使用的评估框架。它在可观测层之上运行,可生成系统、执行轨迹和节点三个层级的文本分析洞察。该框架支持无缝集成,并配备直观用户界面,显著提升评估可及性。在四个基准、七个智能体设置及数万次大模型调用的实验中,Agentic CLEAR展现出高质量、数据驱动的反馈能力,其错误识别与人工标注高度一致,且能有效预测任务成功率。

原文摘要 · Abstract (English)

Agentic systems are becoming more capable: agents define strategies, take actions, and interact with different environments. This autonomy poses serious challenges for overseeing and assessing agent behavior. Most current tools are limited, focusing on observability with basic evaluation capabilities or imposing static, hand-crafted error taxonomies that cannot adapt to new domains. To address this gap, we present Agentic CLEAR, an automatic, dynamic, and easy-to-use evaluation framework. It produces textual insights into the agent behavior on three levels of granularity: system, trace, and node. Agentic CLEAR operates above the observability layer, enabling seamless integration and featuring an intuitive UI that makes agent evaluation highly accessible. In our experiments on four benchmarks, seven agentic settings, and tens of thousands of LLM calls, we show that Agentic CLEAR produces high-quality, data-driven, insightful feedback. Our analysis shows strong alignment with human-annotated errors and the ability to predict task success rate.

智能体评估自动化测试LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。