arXiv:2509.14647cs.AIcs.CL2025-09被引 8

为生产环境中的智能体流程提供可靠评估,发现人类标注遗漏的关键问题。

AgentCompass: Towards Reliable Evaluation of Agentic Workflows in Production

  • 构建多阶段分析流程模拟专家调试思路
  • 在TRAIL基准上达顶尖性能,发现人类忽略的错误
  • 双记忆系统支持持续学习,适合开发团队使用

随着大语言模型在复杂多智能体工作流中的广泛应用,组织面临错误、涌现行为和系统性故障的风险,而现有评估方法难以捕捉。我们提出AgentCompass,首个专为部署后监控与调试设计的评估框架。该框架通过结构化多阶段分析流程——错误识别与分类、主题聚类、量化评分与策略总结——模拟专家调试过程。框架引入情景记忆与语义记忆双重机制,实现执行间的持续学习。通过与设计伙伴合作,在真实部署中验证其实用性,并在公开的TRAIL基准上评估其有效性。AgentCompass在关键指标上达到当前最优表现,同时揭示了人类标注中遗漏的关键问题,凸显其作为开发者友好型工具在生产环境中可靠监控与优化智能体系统的重要作用。

原文摘要 · Abstract (English)

With the growing adoption of Large Language Models (LLMs) in automating complex, multi-agent workflows, organizations face mounting risks from errors, emergent behaviors, and systemic failures that current evaluation methods fail to capture. We present AgentCompass, the first evaluation framework designed specifically for post-deployment monitoring and debugging of agentic workflows. AgentCompass models the reasoning process of expert debuggers through a structured, multi-stage analytical pipeline: error identification and categorization, thematic clustering, quantitative scoring, and strategic summarization. The framework is further enhanced with a dual memory system-episodic and semantic-that enables continual learning across executions. Through collaborations with design partners, we demonstrate the framework's practical utility on real-world deployments, before establishing its efficacy against the publicly available TRAIL benchmark. AgentCompass achieves state-of-the-art results on key metrics, while uncovering critical issues missed in human annotations, underscoring its role as a robust, developer-centric tool for reliable monitoring and improvement of agentic systems in production.

智能体评估LLM应用生产监控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。