调查主流机器学习数据集文档完整性,发现关键信息普遍缺失。
Completeness of Datasets Documentation on ML/AI repositories: an Empirical Investigation
- 构建文档评估框架DTS,明确数据集应包含的核心信息
- 抽查100个热门数据集,发现收集背景与处理流程文档严重不足
- 研究结果警示数据透明度问题,适合关注AI可解释性的研究人员
机器学习/人工智能是过去十年最受关注和资助的计算机科学领域。数据是机器学习的核心,确保使用者充分了解所用数据的质量及其生成过程至关重要,以便追踪、分析并尽可能缓解下游应用中的潜在负面影响。数据集文档是实现这一目标的重要工具。本文旨在实证调查数据集文档实践现状,评估多个流行数据集在机器学习/人工智能仓库中的文档完整性。我们构建了一个数据集文档评估标准——文档检测表(DTS),该标准基于文献中相关研究,明确了数据集应始终附带的关键信息,以保障数据集的合理选择与知情使用。我们使用DTS对来自四个不同仓库的100个热门数据集进行了核查,发现关键信息,特别是数据收集背景与数据处理细节,普遍存在缺失,暴露出严重的不透明性问题。
原文摘要 · Abstract (English)
ML/AI is the field of computer science and computer engineering that arguably received the most attention and funding over the last decade. Data is the key element of ML/AI, so it is becoming increasingly important to ensure that users are fully aware of the quality of the datasets that they use, and of the process generating them, so that possible negative impacts on downstream effects can be tracked, analysed, and, where possible, mitigated. One of the tools that can be useful in this perspective is dataset documentation. The aim of this work is to investigate the state of dataset documentation practices, measuring the completeness of the documentation of several popular datasets in ML/AI repositories. We created a dataset documentation schema -- the Documentation Test Sheet (DTS) -- that identifies the information that should always be attached to a dataset (to ensure proper dataset choice and informed use), according to relevant studies in the literature. We verified 100 popular datasets from four different repositories with the DTS to investigate which information was present. Overall, we observed a lack of relevant documentation, especially about the context of data collection and data processing, highlighting a paucity of transparency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。