arXiv:2606.23515stat.MLcs.LG2026-06

通过贝叶斯实验设计,让数据采集更公平,提升模型公平性与准确率平衡。

FairBED: A Bayesian Experimental Design Approach to Gathering Fairer Data

  • 用敏感属性无关性量化数据公平性,指导数据采集
  • 相比随机采样和传统方法,公平性-准确率权衡显著提升
  • 适合关注数据源头公平性的机器学习研究者

现有机器学习公平性框架多依赖已有数据训练公平模型,但数据本身常含偏见。为此,本文提出FairBED,通过量化数据集对敏感属性的可预测性来评估其公平性,进而构建兼顾目标变量信息增益与敏感属性信息增益最小化的公平感知贝叶斯实验设计(BED)目标。理论证明其与人口均等性相关,实验证明:使用FairBED采集的数据训练出的模型,在公平性与准确率的权衡上优于随机采样和传统BED。该方法从数据采集源头改善公平性。

原文摘要 · Abstract (English)

Frameworks for ensuring fairness in machine learning typically focus on learning fair models from existing data. But this endeavor is often undermined by biases already present in that data. We therefore look to modify the data acquisition process itself to help gather fairer data that is inherently more suitable for training fair predictors. To this end, we introduce FairBED, which provides novel formulations for quantifying the fairness of datasets themselves based on the idea that fair datasets should be uninformative about sensitive attributes. We then use this to construct practical fairness-aware Bayesian experimental design (BED) objectives that maximize expected information gain about the target quantity of interest while minimizing expected information gain about sensitive attributes. We further derive a theoretical link between FairBED and demographic parity, and show empirically that models trained on data gathered using FairBED provide improved fairness-accuracy trade-offs compared to randomly acquired data and conventional BED.

数据公平性贝叶斯设计实验设计公平机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。