arXiv:2505.08638cs.AIcs.CL2025-05被引 82

构建首个系统化评估智能体工作流的公开数据集与分析框架

TRAIL: Trace Reasoning and Agentic Issue Localization

  • 基于真实场景构建148条带标注的智能体工作流轨迹
  • 发现顶级大模型在轨迹调试中准确率仅11%
  • 为智能体错误定位提供可复用的评估标准与工具

随着智能体工作流在多个领域广泛应用,亟需可扩展、系统化的评估方法来分析其生成的复杂执行轨迹。当前方法依赖人工进行长篇轨迹的领域特定分析,难以应对日益增长的复杂性与数据量。由于外部工具输出与语言模型推理之间的相互作用,此类错误分析比传统软件调试更困难。本文提出:(1) 建立稳健动态的智能体工作流轨迹评估需求;(2) 构建智能体系统中常见错误类型的正式分类体系;(3) 提出一个包含148条经人工标注的轨迹数据集TRAIL,该数据集基于该分类体系,并基于成熟的智能体基准构建。为保证生态有效性,数据集涵盖单智能体与多智能体系统,聚焦软件工程与开放世界信息检索等实际应用。评估显示,现代长上下文大模型在轨迹调试任务上表现极差,最佳模型Gemini-2.5-pro仅达到11%准确率。相关数据集与代码已公开,以支持并加速智能体工作流可扩展评估的研究。

原文摘要 · Abstract (English)

The increasing adoption of agentic workflows across diverse domains brings a critical need to scalably and systematically evaluate the complex traces these systems generate. Current evaluation methods depend on manual, domain-specific human analysis of lengthy workflow traces - an approach that does not scale with the growing complexity and volume of agentic outputs. Error analysis in these settings is further complicated by the interplay of external tool outputs and language model reasoning, making it more challenging than traditional software debugging. In this work, we (1) articulate the need for robust and dynamic evaluation methods for agentic workflow traces, (2) introduce a formal taxonomy of error types encountered in agentic systems, and (3) present a set of 148 large human-annotated traces (TRAIL) constructed using this taxonomy and grounded in established agentic benchmarks. To ensure ecological validity, we curate traces from both single and multi-agent systems, focusing on real-world applications such as software engineering and open-world information retrieval. Our evaluations reveal that modern long context LLMs perform poorly at trace debugging, with the best Gemini-2.5-pro model scoring a mere 11% on TRAIL. Our dataset and code are made publicly available to support and accelerate future research in scalable evaluation for agentic workflows.

智能体评估轨迹分析大模型测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。