arXiv:2605.26079cs.CL2026-05被引 4

用智能体自动检测AI评测任务中的漏洞,发现超四分之一存在严重问题。

Automated Benchmark Auditing for AI Agents and Large Language Models

论文配图:Automated Benchmark Auditing for AI Agents and Large Language Models
图 1 · 摘自论文原文
  • 构建智能体框架,自动审查评测任务的设计与执行环境
  • 在168个前沿评测中发现25.7%的任务存在设计缺陷或错误真值
  • 修复问题后模型排名变化显著,性能平均提升9.6%以上

现代AI评测任务复杂度已超出传统验证方法的范围。由领域专家设计的任务常隐含假设、环境说明不全、评估逻辑脆弱,人工标注难以可靠发现。我们提出自动化评测审计框架ABA,系统性审查单个评测任务,揭示隐藏环境依赖、规范缺失和有限评分逻辑等问题。在涵盖九个领域的168个前沿大模型评测及过往NeurIPS论文中应用ABA,发现超过25.7%的任务存在严重问题,包括任务设计模糊、执行环境冲突和错误真值。专家评审与第三方报告验证了审计精度。关键的是,这些问题任务严重扭曲模型能力评估:剔除有缺陷任务后,SWE-bench Verified和Terminal-Bench 2的平均性能分别提升9.9%和9.6%,模型排名亦发生显著变化。我们开源该智能体工具及全部任务标注,以支持未来前沿评测的发展。

原文摘要 · Abstract (English)

Modern AI benchmarks operate at a complexity that outpaces traditional verification methods. Tasks authored by domain experts often contain implicit assumptions, incomplete environment specifications, and brittle evaluation logic that human annotation cannot reliably catch. We introduce Auto Benchmark Audit (ABA), an agentic framework that systematically audits individual benchmark tasks, uncovering issues such as hidden environment dependencies, specification gaps, and limited grading logic. We run ABA on a collection of frontier LLM benchmarks and previous NeurIPS publications, totaling 168 benchmarks across nine domains. Across this corpus, ABA identifies critical issues including ambiguous task design, execution environment conflicts, and incorrect ground truths in over 25.7% of the evaluated tasks. The precision of these automated audits is validated by expert review and independent third-party reports such as upstream PRs. Crucially, we demonstrate that these problematic tasks severely distorts capability assessments for agents and LLMs: filtering out these tasks with issues shifts model rankings and increases average performance on SWE-bench Verified and Terminal-Bench 2 by 9.9% and 9.6%, respectively. We release the agentic tool and all task annotations to support the future development of frontier benchmarks.

评测审计大模型评估智能体基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。