arXiv:2411.12990cs.AIcs.LG2024-11NeurIPS被引 154

评估24个AI基准测试,发现普遍质量差且难复现,提出改进 checklist。

BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices

  • 构建46项最佳实践框架,系统评估基准测试全生命周期
  • 发现多数基准缺乏统计显著性报告与可复现性支持
  • 提供最小质量保障清单和可公开查询的评估仓库

AI模型在高风险场景中日益普及,需对其能力与风险进行充分评估。基准测试是衡量模型性能、追踪进展、识别弱点的重要工具,也影响下游任务选型与政策制定。然而,基准测试质量参差不齐,其设计与可用性直接影响评估效果。本文提出涵盖AI基准测试全生命周期的46项最佳实践框架,对24个主流基准测试进行评估。结果表明,存在显著的质量差异,常用基准普遍存在严重问题:多数未报告结果的统计显著性,且难以复现。为帮助开发者提升质量,我们提供一份最小质量保障检查清单,并建立持续更新的评估仓库(betterbench.stanford.edu),以促进基准测试间的可比性。

原文摘要 · Abstract (English)

AI models are increasingly prevalent in high-stakes environments, necessitating thorough assessment of their capabilities and risks. Benchmarks are popular for measuring these attributes and for comparing model performance, tracking progress, and identifying weaknesses in foundation and non-foundation models. They can inform model selection for downstream tasks and influence policy initiatives. However, not all benchmarks are the same: their quality depends on their design and usability. In this paper, we develop an assessment framework considering 46 best practices across an AI benchmark's lifecycle and evaluate 24 AI benchmarks against it. We find that there exist large quality differences and that commonly used benchmarks suffer from significant issues. We further find that most benchmarks do not report statistical significance of their results nor allow for their results to be easily replicated. To support benchmark developers in aligning with best practices, we provide a checklist for minimum quality assurance based on our assessment. We also develop a living repository of benchmark assessments to support benchmark comparability, accessible at betterbench.stanford.edu.

AI评估基准测试可复现性质量评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。