arXiv:2604.24955cs.CLcs.AI2026-04被引 10

用大模型自动审查大模型评测基准,发现12个致命错误

BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks

论文配图:BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
图 1 · 摘自论文原文
  • 用前沿大模型作为审计员,交叉验证评测任务全流程
  • 在两个科学评测中发现12处作者确认的严重问题,准确率83.3%
  • 每轮审计成本低于15美元,适合大规模评测体系

随着评测基准复杂度提升,许多看似智能体失败的情况实为评测本身缺陷:规范错误、隐含假设与僵化评估脚本误判合理方案。本文提出通过前沿大模型系统性审计评测基础设施,并实现为BenchGuard——首个面向任务导向、执行型智能体评测的自动化审计框架。BenchGuard通过结构化大模型协议交叉验证所有评测资源,可选引入智能体解法或执行轨迹作为诊断依据。部署于两个主流科学评测基准上,BenchGuard在ScienceAgentBench中识别出12项作者确认的问题,包括导致任务不可解的致命错误;在BIXBench Verified-50子集上精准匹配83.3%专家识别问题,捕获了此前人工审查遗漏的缺陷。对50个复杂生物信息学任务的完整审计成本低于15美元,表明自动化评测审计已成为人工审查的有效补充。研究指向一种新型AI辅助评测开发模式:前沿模型不仅是评估对象,更可作为验证评估体系本身的主动参与者。

原文摘要 · Abstract (English)

As benchmarks grow in complexity, many apparent agent failures are not failures of the agent at all - they are failures of the benchmark itself: broken specifications, implicit assumptions, and rigid evaluation scripts that penalize valid alternative approaches. We propose employing frontier LLMs as systematic auditors of evaluation infrastructure, and realize this vision through BenchGuard, the first automated auditing framework for task-oriented, execution-based agent benchmarks. BenchGuard cross-verifies all benchmark artifacts via structured LLM protocols, optionally incorporating agent solutions or execution traces as additional diagnostic evidence. Deployed on two prominent scientific benchmarks, BenchGuard identified 12 author-confirmed issues in ScienceAgentBench - including fatal errors rendering tasks unsolvable - and exactly matched 83.3% of expert-identified issues on the BIXBench Verified-50 subset, catching defects that prior human review missed entirely. A full audit of 50 complex bioinformatics tasks costs under USD 15, making automated benchmark auditing a practical and valuable complement to human review. These findings point toward AI-assisted benchmark development, where frontier models serve not only as subjects of evaluation but as active participants in validating the evaluation infrastructure itself.

大模型评测自动化审计基准验证AI治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。