为前沿AI模型制定基准淘汰标准,防止评估失真。
Deprecating Benchmarks: Criteria and Framework
- 提出基准全/部分淘汰的判断标准
- 建立可操作的基准淘汰框架
- 适合评估者、政策制定者参考
随着前沿人工智能模型快速演进,基准测试在比较不同模型及衡量其在特定任务领域的进展中发挥关键作用。然而,目前缺乏明确指导,说明何时以及如何淘汰已失效的基准。这可能导致评分过高夸大模型能力,甚至掩盖真实能力或进行安全洗白。基于对基准测试实践的回顾,本文提出判断基准应完全或部分淘汰的标准,并构建相应的淘汰框架。研究旨在推动基准测试向更严谨、高质量的方向发展,尤其针对前沿模型,建议对基准开发者、使用者以及政府、学术界和产业界治理机构、政策制定者具有参考价值。
原文摘要 · Abstract (English)
As frontier artificial intelligence (AI) models rapidly advance, benchmarks are integral to comparing different models and measuring their progress in different task-specific domains. However, there is a lack of guidance on when and how benchmarks should be deprecated once they cease to effectively perform their purpose. This risks benchmark scores over-valuing model capabilities, or worse, obscuring capabilities and safety-washing. Based on a review of benchmarking practices, we propose criteria to decide when to fully or partially deprecate benchmarks, and a framework for deprecating benchmarks. Our work aims to advance the state of benchmarking towards rigorous and quality evaluations, especially for frontier models, and our recommendations are aimed to benefit benchmark developers, benchmark users, AI governance actors (across governments, academia, and industry panels), and policy makers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。