arXiv:2606.29159cs.AI2026-06

主流离线故障分析榜单排名误导工程师,因系统差异被掩盖

Pooled Leaderboards Hide System-Specific Winners: A Reporting-Protocol Audit of Offline Root-Cause Analysis Benchmarks

论文配图:Pooled Leaderboards Hide System-Specific Winners: A Reporting-Protocol Audit of Offline Root-Cause Analysis Benchmarks
图 1 · 摘自论文原文
  • 用配对比较法检验11个子系统中的方法表现稳定性
  • 5/6组对比显示子系统效应显著,最高误差达24.8个百分点
  • 提供可复现的审计工具,帮助识别被隐藏的最优方案

离线故障根因分析(RCA)基准常以跨多个子系统的合并top-1准确率排名方法,工程师据此推断适用于自身系统。我们对OpenRCA、RCAEval和PetShop三个公开基准家族进行审计,覆盖11个子系统与778个匹配评分单元。为保证案例一致性,主分析保留四种完整覆盖的方法:BARO、CD-1min适配器、max-|Z|和按服务告警数。六组配对比较均显示子系统层面效应方向不一,所有随机效应95%预测区间包含零值,且6组中有5组在案例级交互检验中拒绝可交换性。剔除一个子系统后选择的最优方法,在最多5个保留子系统上表现更差,最大遗憾达24.8个百分点(RCAEval / Sock-Shop)。我们发布320行审计模块,给定匹配的RCA基准评分表即可重算相同子系统稳定性检查及合并得分。

原文摘要 · Abstract (English)

Offline root-cause-analysis (RCA) benchmarks commonly rank methods by a single pooled top-1 accuracy across multiple subsystems, and engineers often read the pooled winner as a recommendation for their own subsystem. We audit that reading on three public RCA benchmark families -- OpenRCA, RCAEval, and PetShop -- covering 11 subsystems and 778 matched scoring units. To keep pairwise comparisons on identical cases, the main analysis retains four methods or comparators with complete coverage: BARO, a CD-1min adapter, max-$|Z|$, and per-service alert-count. All six pairwise comparisons show subsystem-level effects of both signs, every random-effects 95\% prediction interval crosses zero, and case-level interaction tests reject exchangeability in 5 of 6 pairs. Leave-one-system-out selection picks the lower-scoring method on up to 5 of 11 held-out subsystems, with regret reaching 24.8 pp on RCAEval / Sock-Shop. We release a 320-line audit module; given a matched RCA benchmark score table, it recomputes the same per-subsystem stability checks alongside pooled scores.

故障分析基准测试可解释性评估审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。