arXiv:2409.00252cs.LG2024-09被引 16

来自18位数据集创建者的7条责任设计建议,助你构建更可信的机器学习数据集。

Building Better Datasets: Seven Recommendations for Responsible Design from Dataset Creators

  • 通过访谈18位顶尖数据集创建者,提炼出责任设计的核心方法。
  • 强调数据质量、文档完整性和隐私保护等关键环节。
  • 适合数据集设计者、研究团队及关注伦理AI的开发者参考。

机器学习对高质量数据集的需求日益增长,但其伦理与责任创建仍存隐忧。本文基于对18位领先数据集创建者的定性研究,揭示了他们在实践中面临的挑战与考量。研究发现,当前数据集开发中亟需深化协作、知识共享与集体建设。通过深入分析这些创建者的经验,我们提出七项核心建议,涵盖数据质量、文档规范、隐私与知情同意,以及如何减轻潜在滥用风险。本研究旨在推动对数据集创建过程的深度反思,促进负责任的数据实践,建立对这一关键但常被忽视的研究环节的更全面理解。

原文摘要 · Abstract (English)

The increasing demand for high-quality datasets in machine learning has raised concerns about the ethical and responsible creation of these datasets. Dataset creators play a crucial role in developing responsible practices, yet their perspectives and expertise have not yet been highlighted in the current literature. In this paper, we bridge this gap by presenting insights from a qualitative study that included interviewing 18 leading dataset creators about the current state of the field. We shed light on the challenges and considerations faced by dataset creators, and our findings underscore the potential for deeper collaboration, knowledge sharing, and collective development. Through a close analysis of their perspectives, we share seven central recommendations for improving responsible dataset creation, including issues such as data quality, documentation, privacy and consent, and how to mitigate potential harms from unintended use cases. By fostering critical reflection and sharing the experiences of dataset creators, we aim to promote responsible dataset creation practices and develop a nuanced understanding of this crucial but often undervalued aspect of machine learning research.

数据集设计机器学习伦理责任

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。