arXiv:2507.02825cs.AI2025-07被引 80

提出评估AI智能体的检查清单,解决基准测试中的设计漏洞。

Establishing Best Practices for Building Rigorous Agentic Benchmarks

  • 基于经验与调研制定智能体基准检查清单
  • 在复杂基准上使性能高估降低33%
  • 适合评估智能体能力的研究者和开发者

基准测试对量化追踪人工智能进展至关重要。随着AI智能体能力增强,研究者提出了用于评估复杂现实任务表现的智能体基准。这些基准通常通过特定奖励设计评估任务结果。然而,我们发现许多智能体基准存在任务设置或奖励设计问题:例如,SWE-bench Verified测试用例不足,TAU-bench将空响应计为成功。这些问题可能导致智能体性能被高估或低估达100%相对误差。为提升评估严谨性,我们提出智能体基准检查清单(ABC),该清单综合了我们的构建经验、最佳实践调研及已报告问题。应用于评估设计复杂的CVE-Bench时,ABC将性能过估计减少33%。

原文摘要 · Abstract (English)

Benchmarks are essential for quantitatively tracking progress in AI. As AI agents become increasingly capable, researchers and practitioners have introduced agentic benchmarks to evaluate agents on complex, real-world tasks. These benchmarks typically measure agent capabilities by evaluating task outcomes via specific reward designs. However, we show that many agentic benchmarks have issues in task setup or reward design. For example, SWE-bench Verified uses insufficient test cases, while TAU-bench counts empty responses as successful. Such issues can lead to under- or overestimation of agents' performance by up to 100% in relative terms. To make agentic evaluation rigorous, we introduce the Agentic Benchmark Checklist (ABC), a set of guidelines that we synthesized from our benchmark-building experience, a survey of best practices, and previously reported issues. When applied to CVE-Bench, a benchmark with a particularly complex evaluation design, ABC reduces the performance overestimation by 33%.

智能体评估基准测试可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。