arXiv:2506.13776cs.AIcs.CY2025-06中稿 · ICML被引 3

提升人类基准的严谨性与透明度,让AI性能评估更可信

Recommendations and Reporting Checklist for Rigorous & Transparent Human Baselines in Model Evaluations

  • 提出一套设计、执行和报告人类基准的方法框架
  • 系统审查115项研究,发现现有方法普遍存在不透明问题
  • 提供可复用的检查清单,适合评估者与政策制定者使用

本文主张,基础模型评估中的人类基准必须更加严谨和透明,才能实现人类与AI性能的有意义比较。当前许多研究声称模型达到‘超人水平’,但现有基准方法在严谨性和文档完整性上不足,难以可靠衡量性能差异。基于对测量理论与AI评估文献的元综述,我们构建了一个包含设计、执行与报告建议的框架,并将其整合为一份检查清单。我们利用该清单系统审查了115项基础模型评估中的人类基准研究,识别出普遍存在的方法缺陷;该检查清单亦可指导研究人员开展并报告人类基准实验。我们希望推动更严格的AI评估实践,服务于科研界与政策制定者。数据可在 https://github.com/kevinlwei/human-baselines 获取。

原文摘要 · Abstract (English)

In this position paper, we argue that human baselines in foundation model evaluations must be more rigorous and more transparent to enable meaningful comparisons of human vs. AI performance, and we provide recommendations and a reporting checklist towards this end. Human performance baselines are vital for the machine learning community, downstream users, and policymakers to interpret AI evaluations. Models are often claimed to achieve "super-human" performance, but existing baselining methods are neither sufficiently rigorous nor sufficiently well-documented to robustly measure and assess performance differences. Based on a meta-review of the measurement theory and AI evaluation literatures, we derive a framework with recommendations for designing, executing, and reporting human baselines. We synthesize our recommendations into a checklist that we use to systematically review 115 human baselines (studies) in foundation model evaluations and thus identify shortcomings in existing baselining methods; our checklist can also assist researchers in conducting human baselines and reporting results. We hope our work can advance more rigorous AI evaluation practices that can better serve both the research community and policymakers. Data is available at: https://github.com/kevinlwei/human-baselines

模型评估人类基准透明度可复现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。