arXiv:2608.11323cs.AIcs.LG2026-08

现有智能体排行榜只反映任务专长,不能衡量真实能力。

Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations

  • 用广义可靠性理论分解评估方差,揭示排行榜本质是任务专长排名。
  • 最难任务上评估可靠性降为0,训练表现越好的模型越难复现。
  • 提出部署决策可靠性框架,帮助企业做出可辩护的选型决策。

企业从业者将智能体排行榜视为能力排名,但我们发现,在三个开放追踪基准(TheAgentCompany、$τ^2$-bench 和 AppWorld)中,智能体主效应在所有数据集和检查类型中占比均不足3%,而智能体与任务交互效应占7%-23%。排行榜反映的是任务专长而非通用能力。通过四面广义可靠性理论方差分解,并使用三种估计器(Henderson Method-I、REML via lme4、贝叶斯二项GLMM)验证,结果一致到三位小数。进一步发现:第一,最难题目组的聚合可靠性崩溃,$τ^2$动作检查的$Eρ^2$从0.752降至0.000;第二,训练单元可靠性与保留可靠性负相关($r = -0.90$),即表现越可靠的模型越难复现;第三,群体级诊断在企业基准间可迁移(能力差距比稳定在0.35-0.40),但按家族排序反转;第四,在MAST失败分类中,痕迹级模式特征个体差异大(MAE=0.261),而单元级模式具可泛化性(MAE=0.056,$r=0.83$)。我们据此构建部署决策可靠性(DDR)框架,将方差成分表转化为企业买家可辩护的五个决策。所有代码、数据加载器及拟合产物均已开源。

原文摘要 · Abstract (English)

Enterprise practitioners read agent leaderboards as if they ranked agent capability. We show, across three open agent-trace benchmarks (TheAgentCompany, $τ^2$-bench, and AppWorld), that the agent main effect accounts for less than 3% of total variance in every dataset and check type, while the agent-by-task interaction accounts for 7-23%. Leaderboards rank specialization, not capability. We arrive at this through a four-facet Generalizability Theory variance decomposition, fit with three estimators (Henderson Method-I, REML via lme4, and a Bayesian binomial GLMM) that agree to three decimal places. Four further findings sharpen what the leaderboard is hiding. First, aggregate reliability collapses on the hardest task quartile: $Eρ^2$ on $τ^2$ action_checks falls from 0.752 to 0.000. Second, training-cell reliability negatively correlates with held-out reliability ($r = -0.90$ on $τ^2$), meaning the designs that look most reliable replicate worst. Third, population-level diagnostics transfer across enterprise benchmarks (capability-gap ratio stable at 0.35-0.40) but per-family agent rankings invert. Fourth, on the MAST failure taxonomy, trace-level mode profiles are idiosyncratic (MAE = 0.261) while cell-level profiles generalise (MAE = 0.056, $r = 0.83$). We package these into Deployment Decision Reliability (DDR), a one-page reporting discipline that turns the variance-component table into five decisions an enterprise buyer can defend. All code, data loaders, and fit artifacts are released under an open-source license.

智能体评估可靠性分析企业应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。