arXiv:2510.22363cs.LGcs.CY2025-10被引 1

构建44个公平性数据集与工具包,提升机器学习公平性研究的可复现性。

Bias Begins with Data: The FairGround Corpus for Robust and Reproducible Research on Algorithmic Fairness

  • 整合44个带公平性元信息的表格数据集,统一处理流程
  • 提供标准化加载与预处理工具,支持可复现实验
  • 适合关注算法公平性、可复现研究的研究者使用

随着机器学习系统在高风险决策领域广泛应用,确保其输出公平性成为核心挑战。公平性研究依赖的数据集往往选择范围狭窄、处理不一致且缺乏多样性,严重影响结果的可推广性和可复现性。为此,我们提出FairGround:一个统一框架、数据语料库及Python工具包,旨在推动公平性机器学习研究的可复现性与批判性数据研究。FairGround目前包含44个表格数据集,每个均标注了丰富的公平性相关元信息。配套Python包实现了数据集加载、预处理、转换与划分的标准化,简化实验流程。通过提供多样且文档完善的语料库与强大工具链,FairGround助力开发更公平、可靠且可复现的机器学习模型。所有资源公开可用,支持开放协作研究。

原文摘要 · Abstract (English)

As machine learning (ML) systems are increasingly adopted in high-stakes decision-making domains, ensuring fairness in their outputs has become a central challenge. At the core of fair ML research are the datasets used to investigate bias and develop mitigation strategies. Yet, much of the existing work relies on a narrow selection of datasets--often arbitrarily chosen, inconsistently processed, and lacking in diversity--undermining the generalizability and reproducibility of results. To address these limitations, we present FairGround: a unified framework, data corpus, and Python package aimed at advancing reproducible research and critical data studies in fair ML classification. FairGround currently comprises 44 tabular datasets, each annotated with rich fairness-relevant metadata. Our accompanying Python package standardizes dataset loading, preprocessing, transformation, and splitting, streamlining experimental workflows. By providing a diverse and well-documented dataset corpus along with robust tooling, FairGround enables the development of fairer, more reliable, and more reproducible ML models. All resources are publicly available to support open and collaborative research.

公平性数据集可复现机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。