arXiv:2411.08297cs.LGcs.AI2024-11

不重训练模型,用概率塔性质消除预测中的偏见。

TowerDebias: A Novel Unfairness Removal Method Based on the Tower Property

  • 基于概率论塔性质设计后处理方法,无需了解原模型结构。
  • 在回归与分类任务中显著降低敏感属性影响,提升公平性。
  • 适合部署在黑箱模型上,适用于医疗、金融等敏感场景。

决策过程日益依赖复杂的机器学习工具,引发对预测公平性的关注,尤其涉及种族、性别等敏感群体时。商业使用的“黑箱”模型带来法律与伦理风险。用户交互时,如何削弱敏感属性对预测的影响成为关键挑战。本文提出towerDebias(tDB),一种新型后处理方法,利用概率论中的塔性质,在不重新训练原始模型的前提下,有效降低敏感属性对预测结果的影响。该方法无需了解原模型内部结构,适用范围广。我们给出了tDB的公平性改进形式化定理,并在多个真实数据集上的回归与分类任务中验证其有效性。

原文摘要 · Abstract (English)

Decision-making processes have increasingly come to rely on sophisticated machine learning tools, raising critical concerns about the fairness of their predictions with respect to sensitive groups. The widespread adoption of commercial "black-box" models necessitates careful consideration of their legal and ethical implications for consumers. When users interact with such black-box models, a key challenge arises: how can the influence of sensitive attributes, such as race or gender, be mitigated or removed from its predictions? We propose towerDebias (tDB), a novel post-processing method designed to reduce the influence of sensitive attributes in predictions made by black-box models. Our tDB approach leverages the Tower Property from probability theory to improve prediction fairness without requiring retraining of the original model. This method is highly versatile, as it requires no prior knowledge of the original algorithm's internal structure and is adaptable to a diverse range of applications. We present a formal fairness improvement theorem for tDB and showcase its effectiveness in both regression and classification tasks using multiple real-world datasets.

公平性黑箱模型后处理概率论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。