arXiv:2411.00266cs.LG2024-11NeurIPS综述被引 2

梳理NeurIPS数据集管理实践,揭示数据来源与分发不规范问题。

A Systematic Review of NeurIPS Dataset Management Practices

  • 系统审查NeurIPS数据集的来源、分发、伦理披露和许可信息。
  • 多数数据集来源模糊,托管平台缺乏结构化元数据和版本控制。
  • 呼吁建立标准化数据基础设施,助力学术透明与伦理合规。

随着机器学习方法对更大训练数据集的需求,研究人员在数据管理方面面临重大挑战。尽管已建立伦理审查、文档规范和检查清单,但社区内是否存在一致的数据管理实践仍不明确。这种缺乏全面概览的情况阻碍了我们诊断和解决大规模数据管理中的根本矛盾与伦理问题。本文对发表于NeurIPS数据集与基准测试专题的数据集进行了系统性审查,重点关注四个关键方面:数据来源、分发方式、伦理披露和许可协议。研究发现,由于过滤与整理过程不清晰,数据来源往往不明;虽然使用了多种托管平台,但仅有少数提供结构化元数据和版本控制。这些不一致性凸显了建立标准化数据基础设施以支持数据发布与管理的迫切需求。

原文摘要 · Abstract (English)

As new machine learning methods demand larger training datasets, researchers and developers face significant challenges in dataset management. Although ethics reviews, documentation, and checklists have been established, it remains uncertain whether consistent dataset management practices exist across the community. This lack of a comprehensive overview hinders our ability to diagnose and address fundamental tensions and ethical issues related to managing large datasets. We present a systematic review of datasets published at the NeurIPS Datasets and Benchmarks track, focusing on four key aspects: provenance, distribution, ethical disclosure, and licensing. Our findings reveal that dataset provenance is often unclear due to ambiguous filtering and curation processes. Additionally, a variety of sites are used for dataset hosting, but only a few offer structured metadata and version control. These inconsistencies underscore the urgent need for standardized data infrastructures for the publication and management of datasets.

数据管理机器学习伦理合规标准化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。