提出12项指标评估AI代理的可靠性,揭示其实际表现与测试成绩的差距。
Towards a Science of AI Agent Reliability
- 从一致性、鲁棒性、可预测性和安全性四维度构建可靠性评估体系
- 15个模型在双基准测试中显示能力提升但可靠性改善有限
- 为开发者提供分析代理失效机制的新工具,适合关注AI安全的团队
AI代理被越来越多地用于执行关键任务。尽管标准基准上的准确率持续上升,许多代理在实际应用中仍频繁失败。这种差异凸显了当前评估方式的根本局限:将代理行为压缩为单一成功指标,掩盖了其在多次运行中的一致性、对扰动的抵抗能力、故障的可预测性以及错误严重程度的边界等关键问题。基于安全关键工程,我们提出12项具体指标,从一致性、鲁棒性、可预测性和安全性四个维度全面刻画代理可靠性。在两个互补基准上评估15个模型,发现近期能力提升仅带来可靠性的小幅改善。本研究通过暴露这些持久性局限,补充传统评估方法,并为理解代理如何表现、退化和失效提供了工具。
原文摘要 · Abstract (English)
AI agents are increasingly deployed to execute important tasks. While rising accuracy scores on standard benchmarks suggest rapid progress, many agents still continue to fail in practice. This discrepancy highlights a fundamental limitation of current evaluations: compressing agent behavior into a single success metric obscures critical operational flaws. Notably, it ignores whether agents behave consistently across runs, withstand perturbations, fail predictably, or have bounded error severity. Grounded in safety-critical engineering, we provide a holistic performance profile by proposing twelve concrete metrics that decompose agent reliability along four key dimensions: consistency, robustness, predictability, and safety. Evaluating 15 models across two complementary benchmarks, we find that recent capability gains have only yielded small improvements in reliability. By exposing these persistent limitations, our metrics complement traditional evaluations while offering tools for reasoning about how agents perform, degrade, and fail.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。