arXiv:2511.07032cs.LGstat.ML2025-11AAAI被引 1

通过数据选择提升模型公平性,无需显式约束。

Fair Bayesian Data Selection via Generalized Discrepancy Measures

  • 基于贝叶斯框架,用分布差异度量对齐各组数据后验分布。
  • 在多个基准数据集上,公平性和准确率均优于现有方法。
  • 适合关注数据偏见修复的从业者,尤其强调可扩展性。

随着机器学习在高风险场景中的应用增多,公平性问题日益突出。现有公平性方法多在模型层面干预,但普遍存在计算成本高、可扩展性差和泛化能力弱的问题。为此,我们提出一种贝叶斯数据选择框架,通过将各组模型参数与样本权重的后验分布与共享中心分布对齐,实现公平性保障。该框架支持多种分布差异度量(如Wasserstein距离、最大均值差异、f-散度),实现几何感知的灵活对齐,无需施加显式公平约束。此数据驱动方法有效缓解训练数据中的群体偏差,提升下游任务的公平性,并具备理论保证。在多个基准数据集上的实验表明,本方法在公平性和准确性方面均持续优于现有的数据选择与模型级公平方法。

原文摘要 · Abstract (English)

Fairness concerns are increasingly critical as machine learning models are deployed in high-stakes applications. While existing fairness-aware methods typically intervene at the model level, they often suffer from high computational costs, limited scalability, and poor generalization. To address these challenges, we propose a Bayesian data selection framework that ensures fairness by aligning group-specific posterior distributions of model parameters and sample weights with a shared central distribution. Our framework supports flexible alignment via various distributional discrepancy measures, including Wasserstein distance, maximum mean discrepancy, and $f$-divergence, allowing geometry-aware control without imposing explicit fairness constraints. This data-centric approach mitigates group-specific biases in training data and improves fairness in downstream tasks, with theoretical guarantees. Experiments on benchmark datasets show that our method consistently outperforms existing data selection and model-based fairness methods in both fairness and accuracy.

公平性贝叶斯方法数据选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。