提出生产环境智能体评估框架,揭示七类真实故障并解决指标失效问题。
Evaluating Agentic AI in the Wild: Failure Modes, Drift Patterns, and a Production Evaluation Framework
- 构建七类生产级智能体故障分类体系,基于百亿事件运行数据
- 标准指标在四类故障上完全无效,三类滞后多个周期才暴露
- 推出可连续评估的PAEF框架,开源实现支持真实流量监控
现有大模型评估框架(如HELM、MT-Bench、AgentBench、BIG-bench)均针对受控、单次会话、实验室规模场景设计,无法应对智能体系统在生产环境中持续运行时出现的挑战:决策错误累积、工具失败级联、输出非确定性漂移,以及长程任务缺乏真实标签。本文有三项贡献:第一,基于百亿事件规模系统的观测,提出七类独特的生产级智能体故障模式分类;第二,实证证明主流指标(ROUGE、BERTScore、准确率/AUC及上述代理基准)无法检测其中四类故障,其余三类仅在多次评估周期后才显现;第三,提出PAEF(生产级智能体评估框架),包含五个维度的评估体系,并提供开源参考实现,专为持续评估生产流量而非离散基准测试设计。分析表明,标准指标对四类故障完全无感,另三类需延迟多个周期才能捕捉。
原文摘要 · Abstract (English)
Existing evaluation frameworks for large language models -- including HELM, MT-Bench, AgentBench, and BIG-bench -- are designed for controlled, single-session, lab-scale settings. They do not address the evaluation challenges that emerge when agentic AI systems operate continuously in production: compounding decision errors, tool failure cascades, non-deterministic output drift, and the absence of ground truth for long-horizon tasks. This paper makes three contributions. First, we present a taxonomy of seven failure modes unique to production agentic systems, each grounded in observations from systems operating at billion-event scale. Second, we demonstrate empirically where standard metrics -- ROUGE, BERTScore, accuracy/AUC, and the agentic benchmarks above -- fail to detect each failure mode. Third, we propose PAEF (Production Agentic Evaluation Framework), a five-dimension evaluation framework with an open-source reference implementation, designed for continuous evaluation on production traffic rather than episodic benchmark runs. Our analysis shows that standard metrics fail to detect four of the seven failure modes entirely and detect three others only after a lag of multiple evaluation cycles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。