提出可微分公平性损失,兼顾模型性能与公平性。
Differential Adjusted Parity for Learning Fair Representations
- 用可微分的调整均等度量构建统一优化目标
- 在多个数据集上提升公平性指标最高达44.1%
- 避免模型在所有敏感群体上表现均差的退化问题
开发公平且无偏的机器学习模型仍是人工智能领域的关键目标。本文提出可微分调整均等度(Differential Adjusted Parity, DAP)损失,用于生成无偏且信息丰富的表示。该方法采用可微分的调整均等度量,构建统一的目标函数,结合下游任务分类准确率及其在敏感特征域间的不一致性,实现性能提升与偏差缓解的统一优化。核心在于使用软平衡准确率,相较于以往非对抗性方法,DAP避免了因在所有敏感群体上表现均差而满足度量的退化问题。实验表明,DAP在下游任务准确率与公平性方面优于多个对抗性模型:在人口统计均等性、等几率准则和敏感特征准确率上,分别比最优对抗模型提升22.5%、44.1%和40.1%。总体而言,DAP损失及其关联度量对构建更公平的机器学习模型具有重要意义。
原文摘要 · Abstract (English)
The development of fair and unbiased machine learning models remains an ongoing objective for researchers in the field of artificial intelligence. We introduce the Differential Adjusted Parity (DAP) loss to produce unbiased informative representations. It utilises a differentiable variant of the adjusted parity metric to create a unified objective function. By combining downstream task classification accuracy and its inconsistency across sensitive feature domains, it provides a single tool to increase performance and mitigate bias. A key element in this approach is the use of soft balanced accuracies. In contrast to previous non-adversarial approaches, DAP does not suffer a degeneracy where the metric is satisfied by performing equally poorly across all sensitive domains. It outperforms several adversarial models on downstream task accuracy and fairness in our analysis. Specifically, it improves the demographic parity, equalized odds and sensitive feature accuracy by as much as 22.5\%, 44.1\% and 40.1\%, respectively, when compared to the best performing adversarial approaches on these metrics. Overall, the DAP loss and its associated metric can play a significant role in creating more fair machine learning models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。