用AI自动检测智能体评测数据中的漏洞,提升评测可信度。
Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks

- 开发四类缺陷检测器:真相获取、工具失败、猜测漏洞、答案格式模糊。
- 在五大数据集上发现多个已验证质量问题,部分难被人工随机发现。
- 适合评测构建者与研究者使用,推动自动化质量保障落地。
前沿模型的能力常通过智能体评测基准进行评估。为确保结果可信,评测需准确衡量其宣称能力且无致命缺陷。此前对SWE-Bench-Verified等基准的手动审计已揭示多种有效性问题。然而,人工审核难以规模化,且尚不清楚自动化方法能否可靠发现破坏评测有效性的缺陷。本文开发了AI扫描器,用于识别四类有效性问题:真相访问、工具失败、猜测漏洞、答案格式模糊。我们制定了评分标准供人工标注,并在Inspect Evals基准的独立测试集上以人工标签评估扫描器性能。扫描器在五个广泛应用的基准中发现了多个经验证的质量问题,部分问题极难通过随机人工检查发现。但并非所有案例均被识别,扫描器在不同基准、标准和模型上的表现差异显著。我们指出若干开放挑战,包括评估领域标准化不足导致的性能下降。这些结果证明,自动化转录分析可用于更广泛地审计评测质量。
原文摘要 · Abstract (English)
Capabilities of frontier models are often assessed using agentic benchmarks. To trust these results, benchmarks must accurately measure what they claim to and be free from invalidating flaws. Previous manual audits of benchmarks such as SWE-Bench-Verified have uncovered several validity issues in transcripts. However, manual review is difficult to scale, and it is unclear whether automated methods can reliably surface flaws that compromise benchmark validity. In this paper, we developed AI scanners to detect four types of validity issues: ground truth access, tool failure, guessing vulnerability, and answer format ambiguity. We produced grading rubrics for each to instruct human labeling, and evaluated the scanners against human labels on a held-out test set of Inspect Evals benchmarks. Our scanners identified several verified quality issues in five widely used benchmarks, including cases unlikely to be caught by random manual inspection. Not all cases were identified, and scanner performance varied substantially across benchmarks, criteria and models. We highlight several open challenges to be addressed to improve scanners for stronger quality assurance claims, including broader standardization gaps in the evaluation field that degrade scanner performance. Together, these results serve as a proof of concept for using automated transcript analysis to audit benchmark quality more broadly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。