剖析基准数据仓库现状,提出改进机器学习评测的系统性方案
Benchmark Data Repositories for Better Benchmarking
- 系统分析基准数据仓库的设计与使用问题
- 指出数据代表性不足与评估方式单一等关键缺陷
- 适合关注评测规范与数据治理的研究者参考
在机器学习研究中,通常通过算法在标准基准数据集上的表现来评估其性能。尽管已有大量工作对机器学习中的数据与评测实践提出了指导和批评,但针对这些数据集所存储、记录和共享的数据仓库的关注却相对较少。本文分析了‘基准数据仓库’的现状及其在改善评测中的作用,涵盖数据本身的问题(如代表性偏差、构念效度)以及评估方式的问题(如过度依赖少数数据集和指标、缺乏可复现性)。为此,本文识别并讨论了设计与使用基准数据仓库的一系列考量因素,旨在推动机器学习评测实践的改进。
原文摘要 · Abstract (English)
In machine learning research, it is common to evaluate algorithms via their performance on standard benchmark datasets. While a growing body of work establishes guidelines for -- and levies criticisms at -- data and benchmarking practices in machine learning, comparatively less attention has been paid to the data repositories where these datasets are stored, documented, and shared. In this paper, we analyze the landscape of these $\textit{benchmark data repositories}$ and the role they can play in improving benchmarking. This role includes addressing issues with both datasets themselves (e.g., representational harms, construct validity) and the manner in which evaluation is carried out using such datasets (e.g., overemphasis on a few datasets and metrics, lack of reproducibility). To this end, we identify and discuss a set of considerations surrounding the design and use of benchmark data repositories, with a focus on improving benchmarking practices in machine learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。