arXiv:2412.20377cs.LGcs.CY2024-12TPAMI被引 3

揭示数据分布差异如何影响模型公平性,并提出有效提升公平性的训练方法。

On Demographic Group Fairness Guarantees in Deep Learning

  • 建立理论框架,量化不同群体间数据分布差异对公平性的影响。
  • 实验表明种族差异导致的特征分布不均显著影响模型公平性。
  • 提出FAR正则化方法,提升各群体性能与整体公平性指标。

我们提出一个理论框架,分析数据分布与深度学习中公平性保障之间的关系。建立了新的边界,考虑了不同人口群体间分布异质性,推导出公平性误差和收敛速率边界,刻画了分布差异如何影响公平性与准确率的权衡。在多个模态数据集(包括FairVision、CheXpert、HAM10000、FairFace、ACS Income、CivilComments-WILDS)上进行大量实验,验证了理论发现,表明不同群体间的特征分布差异显著影响模型公平性,尤其在种族类别上表现突出。基于这些洞察,我们提出公平感知正则化(FAR),一种最小化群体间特征中心与协方差差异的实用训练目标。FAR在所有数据集上一致提升了总体AUC、ES-AUC及子群体表现。本工作深化了对人工智能系统公平性的理论理解,并为开发更公平的算法提供了基础。分析代码已公开于https://github.com/Harvard-AI-and-Robotics-Lab/FairnessGuarantee。

原文摘要 · Abstract (English)

We present a theoretical framework analyzing the relationship between data distributions and fairness guarantees in equitable deep learning. We establish novel bounds that account for distribution heterogeneity across demographic groups, deriving fairness error and convergence rate bounds that characterize how distributional differences affect the fairness-accuracy trade-off. Extensive experiments across diverse modalities, including FairVision, CheXpert, HAM10000, FairFace, ACS Income, and CivilComments-WILDS, validate our theoretical findings, demonstrating that feature distribution differences across demographic groups significantly impact model fairness, with disparities particularly pronounced in racial categories. Motivated by these insights, we propose Fairness-Aware Regularization (FAR), a practical training objective that minimizes inter-group discrepancies in feature centroids and covariances. FAR consistently improves overall AUC, ES-AUC, and subgroup performance across all datasets. Our work advances the theoretical understanding of fairness in AI systems and provides a foundation for developing more equitable algorithms. The code for analysis is publicly available at https://github.com/Harvard-AI-and-Robotics-Lab/FairnessGuarantee.

公平性深度学习正则化数据分布

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。