arXiv:2605.04624cs.AIcs.SE2026-05

发现智能体修复评估中因评价器信号干扰导致排名不稳定,发布数据集可复现并修复此问题。

AuditRepairBench: A Paired-Execution Trace Corpus for Evaluator-Channel Ranking Instability in Agent Repair

论文配图:AuditRepairBench: A Paired-Execution Trace Corpus for Evaluator-Channel Ranking Instability in Agent Repair
图 1 · 摘自论文原文
  • 通过四类筛选机制构建评分稳定性检测框架,识别评价器信号对修复排序的影响。
  • 实测显示屏蔽特定通道可减少62%的排名波动,远优于随机屏蔽或重训练。
  • 适合关注AI评估可靠性、自动化系统鲁棒性的研究者和开发者使用。

智能体修复排行榜在评价器配置变更时会发生显著重排,其中部分由依赖评价器信号选择修复方案的方法引发。本文在公开排行榜上记录该故障模式,并发布AuditRepairBench,一个包含57.6万注册单元(9.6万已执行)的成对执行轨迹数据集,可在明确定义的可观测边界内操作化评价器通道阻塞引起的排名不稳定性。采用模块化筛选架构,集成四种可替换实现:学习型影响代理、基于规则的通道暴露比(无需训练模型)、反事实敏感性代理、稀疏人工审计代理,共同生成筛选后验分布,驱动单元级翻转函数、集合标签、分层系统评分与集合榜单。该资源经80例源级通道手术子集的机制锚定验证,在独立发现协议下,两组标注员在不知晓筛选设计的情况下盲查出耦合模式,冻结集成在79例中达到0.83的联合AUROC;具备实施稳健性、不确定性传播使95%覆盖率从0.81提升至0.95,以及前向迁移能力(社区评价器斯皮尔曼ρ=0.65)。筛选引导的盲区修补在少于50行代码下将排名位移降低55%~74%(均值62%),而随机通道屏蔽最多仅降7%,通用重训练最多降13%。AuditRepairBench-Lite(仅规则配置,1.2万单元子集)在24个GPU小时下维持肯德尔τ=0.88的排行榜一致性,为主要发布成果,大小为42 GB。

原文摘要 · Abstract (English)

Agent-repair leaderboards reorder under evaluator reconfiguration, and a measurable share of the reordering is produced by methods that consult evaluator-derived signal during internal selection of candidate repairs. We document this failure mode on a public leaderboard and release AuditRepairBench, a paired-execution trace corpus of 576,000 registered cells (96,000 executed) that operationalizes evaluator-channel-blocking ranking instability within a declared observability boundary. A modular screening architecture decides pathway-blocking through four interchangeable implementations, a learned influence proxy, a rule-based channel-exposure ratio that uses no trained model, a counterfactual sensitivity proxy, and a sparse human-audit proxy, combined into a screening posterior that feeds a cell-level flip functional, a set-valued label, a stratified system score, and a set-valued leaderboard. The resource is supported by mechanism-anchored validation on an 80-case source-level channel-surgery subset, an independent-discovery protocol under which two annotator groups separated from the pipeline developers discover coupling patterns blinded to the screening design and the frozen ensemble attains pooled AUROC 0.83 on their 79 cases, implementation robustness, uncertainty propagation that raises 95% coverage from 0.81 to 0.95, and forward transfer with pooled community-evaluator Spearman \r{ho} = 0.65. Screening-guided blinding patches reduce rank displacement by 55--74% (mean 62%) at fewer than 50 lines of code, whereas random channel blinding produces at most 7% reduction and generic retraining at most 13%. AuditRepairBench-Lite, a rule-only configuration on a 12,000-cell subset, preserves the leaderboard at Kendall τ = 0.88 under twenty-four GPU-hours and is the primary release artifact at 42 GB.

评估可靠性智能体修复排行榜稳定数据集发布

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。