arXiv:2605.10516cs.AI2026-05被引 2

提出可检验的AI代理可靠性评估方法,区分能力与执行鲁棒性。

Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability

论文配图:Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability
图 1 · 摘自论文原文
  • 用U统计量和核方法量化输出与轨迹一致性
  • 轨迹级指标诊断敏感度远超传统通过率
  • 适合高风险场景下代理系统的缺陷定位

本文建立了严格的测量科学框架,用于量化智能体在语义保持扰动下的可靠性。通过采用U统计量评估输出层面的可靠性,以及基于核的方法衡量轨迹层面的稳定性,提出了一种在多种运行条件下评估智能体的合理方法。研究揭示了智能体核心能力与执行鲁棒性之间的关键差异:即使具备完成任务所需知识,微小的任务级变化也可能导致策略完全崩溃。在三个代理基准上的大量实验验证表明,轨迹级一致性指标的诊断灵敏度显著高于传统的pass@1指标。该框架提供了数学工具,可定位并识别导致部署受阻的架构缺陷,助力高风险真实环境中的智能体可靠应用。

原文摘要 · Abstract (English)

This paper establishes a rigorous measurement science for AI agent reliability, providing a foundational framework for quantifying consistency under semantically preserving perturbations. By leveraging $U$-statistics for output-level reliability and kernel-based metrics for trajectory-level stability, we offer a principled approach to evaluating agents across diverse operating conditions. Our proposal highlights the important distinction between the core capability and execution robustness of an agent, showing that minor task-level variations can induce complete strategy breakdowns despite the agent possessing the requisite knowledge for the task. We validate our framework through extensive experiments on three agentic benchmarks, demonstrating that trajectory-level consistency metrics provide far greater diagnostic sensitivity than traditional pass@1 rates. By providing the mathematical tools to isolate where and why agents deviate, we enable the identification and rectification of architectural concerns that hinder the deployment of agents in high-stakes, real-world environments.

AI可靠性智能体评估统计方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。