审视当前AI评测的可信度,揭示其隐藏缺陷与系统性风险。
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
- 综合100项研究,剖析评测设计与应用中的细粒问题。
- 发现数据偏见、结果可操纵、评估逻辑过时等关键缺陷。
- 适合关注AI安全、评测公正性的研究者与政策制定者阅读。
量化人工智能(AI)评测已成为评估模型性能、能力与安全性的核心工具,深刻影响着AI发展路径,并日益融入监管框架。然而,随着其影响力扩大,对评测在高影响能力、安全性及系统性风险等敏感议题上的有效性与后果的担忧也随之上升。本文对近十年发表的约100项探讨量化评测缺陷的研究进行了跨学科元综述,整合了数据集构建偏见、文档缺失、数据污染、信号噪声混淆等细粒度问题,以及过度聚焦单次文本模型测试、忽视多模态交互与人机协同等更广泛的社技术问题。研究还揭示了当前评测实践中的系统性弊病,如激励错位、构念效度不足、未知未知、结果被“刷分”等。同时指出,评测体系深受文化、商业与竞争动态塑造,常以追求前沿性能为先,牺牲更广泛社会关切。通过呈现现有评测流程的风险,本文质疑对评测的过度信任,推动提升量化评测在真实复杂场景中的问责性与相关性。
原文摘要 · Abstract (English)
Quantitative Artificial Intelligence (AI) Benchmarks have emerged as fundamental tools for evaluating the performance, capability, and safety of AI models and systems. Currently, they shape the direction of AI development and are playing an increasingly prominent role in regulatory frameworks. As their influence grows, however, so too does concerns about how and with what effects they evaluate highly sensitive topics such as capabilities, including high-impact capabilities, safety and systemic risks. This paper presents an interdisciplinary meta-review of about 100 studies that discuss shortcomings in quantitative benchmarking practices, published in the last 10 years. It brings together many fine-grained issues in the design and application of benchmarks (such as biases in dataset creation, inadequate documentation, data contamination, and failures to distinguish signal from noise) with broader sociotechnical issues (such as an over-focus on evaluating text-based AI models according to one-time testing logic that fails to account for how AI models are increasingly multimodal and interact with humans and other technical systems). Our review also highlights a series of systemic flaws in current benchmarking practices, such as misaligned incentives, construct validity issues, unknown unknowns, and problems with the gaming of benchmark results. Furthermore, it underscores how benchmark practices are fundamentally shaped by cultural, commercial and competitive dynamics that often prioritise state-of-the-art performance at the expense of broader societal concerns. By providing an overview of risks associated with existing benchmarking procedures, we problematise disproportionate trust placed in benchmarks and contribute to ongoing efforts to improve the accountability and relevance of quantitative AI benchmarks within the complexities of real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。