arXiv:2504.01223cs.LGmath.PR2025-04

通过分布公平性约束,实现模型后处理的可解释偏见缓解。

Explainable post-training bias mitigation with distribution-based fairness metrics

  • 基于随机梯度下降的后处理框架,适配多种模型类型。
  • 在多个数据集上优于贝叶斯搜索、最优传输等方法。
  • 支持不同公平水平,适合需要透明解释的场景。

我们提出一种新型偏见缓解框架,采用基于分布的公平性约束,适用于生成跨多种公平水平的性别盲和可解释机器学习模型。该框架通过后处理实现,无需重新训练底层模型即可高效生成更公平的模型。基于随机梯度下降,该框架可应用于多种模型类型,尤其针对梯度提升决策树的后处理。此外,我们设计了一类广义全局公平性度量,并开发了与框架兼容的可微且一致的估计器。我们在多个数据集上对方法进行实证测试,并与贝叶斯搜索、最优传输投影及直接神经网络训练等替代方案进行比较。

原文摘要 · Abstract (English)

We develop a novel bias mitigation framework with distribution-based fairness constraints suitable for producing demographically blind and explainable machine-learning models across a wide range of fairness levels. This is accomplished through post-processing, allowing fairer models to be generated efficiently without retraining the underlying model. Our framework, which is based on stochastic gradient descent, can be applied to a wide range of model types, with a particular emphasis on the post-processing of gradient-boosted decision trees. Additionally, we design a broad family of global fairness metrics, along with differentiable and consistent estimators compatible with our framework, building on previous work. We empirically test our methodology on a variety of datasets and compare it with alternative post-processing approaches, including Bayesian search, optimal transport projection, and direct neural network training.

偏见缓解可解释性后处理公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。