滥用数据仓库导致研究质量下降,需警惕模型选择与预处理陷阱
Unreflected Use of Tabular Data Repositories Can Undermine Research Quality
- 分析公开数据集使用中的三类常见问题:模型选择不当、忽略强基线、预处理错误
- 以OpenML为例,揭示多个顶会研究因依赖数据仓库而出现方法缺陷
- 建议改进数据仓库设计,提升研究可复现性与科学严谨性
数据仓库已积累大量来自不同领域的表格数据集,机器学习研究者广泛利用这些数据集评估新方法。因此,数据仓库在表格数据研究中具有重要地位,不仅托管数据,还提供监督学习任务的使用说明。本文指出,尽管可用性已有显著提升,但对数据仓库的无反思使用可能已削弱研究质量与科学严谨性。我们通过近期若干知名研究案例,展示其在使用OpenML等大型表格数据仓库时存在的问题,包括(1)采用次优的模型选择策略,(2)忽视强大基线模型,(3)不恰当的数据预处理。这些案例揭示了潜在风险。针对此,我们讨论数据仓库如何通过机制改进,防止数据误用,从而成为提升实证研究整体质量的基石。
原文摘要 · Abstract (English)
Data repositories have accumulated a large number of tabular datasets from various domains. Machine Learning researchers are actively using these datasets to evaluate novel approaches. Consequently, data repositories have an important standing in tabular data research. They not only host datasets but also provide information on how to use them in supervised learning tasks. In this paper, we argue that, despite great achievements in usability, the unreflected usage of datasets from data repositories may have led to reduced research quality and scientific rigor. We present examples from prominent recent studies that illustrate the problematic use of datasets from OpenML, a large data repository for tabular data. Our illustrations help users of data repositories avoid falling into the traps of (1) using suboptimal model selection strategies, (2) overlooking strong baselines, and (3) inappropriate preprocessing. In response, we discuss possible solutions for how data repositories can prevent the inappropriate use of datasets and become the cornerstones for improved overall quality of empirical research studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。