深度集成模型对不同群体的性能提升不均,可能加剧算法不公平。
The Disparate Benefits of Deep Ensembles
- 分析深度集成在不同社会群体中的预测表现差异
- 发现集成模型对某些群体的性能提升显著高于其他群体
- 提出用后处理方法可有效缓解这种不公平现象
深度神经网络集成(Deep Ensembles)被广泛用于提升预测性能,但其对算法公平性的影响尚不明确。本文研究了集成性能增益与公平性的关系,发现其对不同社会相关群体(如年龄、性别、种族)的提升存在显著差异,这一现象称为‘非均衡受益效应’。我们在多个包含受保护属性的人脸分析和医学影像数据集上进行实证分析,结果表明该效应影响统计均等性和等机会等多个公平性指标。进一步发现,集成成员在各群体间的预测多样性差异可解释该现象。最后,我们验证了经典的Hardt后处理方法能有效缓解此问题,因其可利用集成模型更优的预测分布校准能力。
原文摘要 · Abstract (English)
Ensembles of Deep Neural Networks, Deep Ensembles, are widely used as a simple way to boost predictive performance. However, their impact on algorithmic fairness is not well understood yet. Algorithmic fairness examines how a model's performance varies across socially relevant groups defined by protected attributes such as age, gender, or race. In this work, we explore the interplay between the performance gains from Deep Ensembles and fairness. Our analysis reveals that they unevenly favor different groups, a phenomenon that we term the disparate benefits effect. We empirically investigate this effect using popular facial analysis and medical imaging datasets with protected group attributes and find that it affects multiple established group fairness metrics, including statistical parity and equal opportunity. Furthermore, we identify that the per-group differences in predictive diversity of ensemble members can explain this effect. Finally, we demonstrate that the classical Hardt post-processing method is particularly effective at mitigating the disparate benefits effect of Deep Ensembles by leveraging their better-calibrated predictive distributions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。