提出可观察性框架,用运行日志分析智能体系统行为,突破传统黑箱评测局限。
Beyond Black-Box Benchmarking: Observability, Analytics, and Optimization of Agentic Systems
- 基于运行日志构建可观测性分析框架,识别动态执行流与异常
- 用户研究显示79%认为非确定性流程是主要挑战,验证问题严重性
- 适合开发与运维人员评估智能体系统可靠性,推动可解释性改进
智能体系统通过协作完成多样化任务,其非确定性、上下文敏感和动态特性给观测、分析与优化带来新挑战。传统评测方法难以应对这些复杂性。本文探讨了在开发、测试和维护阶段分析与优化智能体系统的关键问题,如自然语言变异性与不可预测的执行流,这些因素削弱了系统的可预测性与可控性,需采用自适应策略应对输入变化与行为演化。用户研究表明,79%的受访者认为非确定性执行流是主要挑战。本文通过实证验证主张,强调必须超越传统基准评测。为此,我们提出分类体系,明确预期分析结果及采集方式,扩展标准可观测性框架。在此基础上,提出一种新型评测方法:以智能体运行日志为输入,生成包括发现的执行流与问题在内的分析结果。该方法克服了现有方法的局限,为更全面、智能的评估策略奠定基础,有望促进更具适应性、可解释性和鲁棒性的智能体系统发展。
原文摘要 · Abstract (English)
The rise of agentic AI systems, where agents collaborate to perform diverse tasks, poses new challenges with observing, analyzing and optimizing their behavior. Traditional evaluation and benchmarking approaches struggle to handle the non-deterministic, context-sensitive, and dynamic nature of these systems. This paper explores key challenges and opportunities in analyzing and optimizing agentic systems across development, testing, and maintenance. We explore critical issues such as natural language variability and unpredictable execution flows, which hinder predictability and control, demanding adaptive strategies to manage input variability and evolving behaviors. Through our user study, we supported these hypotheses. In particular, we showed a 79% agreement that non deterministic flow of agentic systems acts as a major challenge. Finally, we validated our statements empirically advocating the need for moving beyond classical benchmarking. To bridge these gaps, we introduce taxonomies to present expected analytics outcomes and the ways to collect them by extending standard observability frameworks. Building on these foundations, we introduce and demonstrate novel approach for benchmarking of agent evaluation systems. Unlike traditional "black box" performance evaluation approaches, our benchmark is built from agent runtime logs as input, and analytics outcome including discovered flows and issues. By addressing key limitations in existing methodologies, we aim to set the stage for more advanced and holistic evaluation strategies, which could foster the development of adaptive, interpretable, and robust agentic AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。