无需人口统计信息,通过谱不确定集实现更可靠的公平性保障
Safe Fairness Guarantees Without Demographics in Classification: Spectral Uncertainty Set Perspective
- 基于傅里叶特征谱调整构建不确定集,约束最差分布偏离实测分布程度
- 在20个州数据上平均公平性指标最优,且结果波动最小
- 适用于无法获取群体标签的场景,尤其适合高风险决策系统
随着自动化分类系统广泛应用,其可能加剧社会偏见的问题日益突出。现有多数方法依赖每个样本的群体信息,但实际中常不可得。无群体信息的公平性方法多采用鲁棒优化,针对一组可能分布(不确定集)进行最坏情况设计。然而,现有方法常过度关注异常值或过度悲观情景,影响整体性能与公平性。为此,本文提出SPECTRE,一种极小极大公平方法:通过调整简单傅里叶特征映射的谱结构,并限制最坏分布相对于经验分布的偏离程度。在涵盖20个州的美国社区调查数据集上进行了广泛实验。SPECTRE在公平性保障方面表现出最高平均值和最小四分位距,优于当前最先进的方法,甚至超过部分需群体信息的方法。此外,理论分析给出了个体群体与总体最坏误差的可计算边界,并刻画了导致极端表现的最坏分布。
原文摘要 · Abstract (English)
As automated classification systems become increasingly prevalent, concerns have emerged over their potential to reinforce and amplify existing societal biases. In the light of this issue, many methods have been proposed to enhance the fairness guarantees of classifiers. Most of the existing interventions assume access to group information for all instances, a requirement rarely met in practice. Fairness without access to demographic information has often been approached through robust optimization techniques,which target worst-case outcomes over a set of plausible distributions known as the uncertainty set. However, their effectiveness is strongly influenced by the chosen uncertainty set. In fact, existing approaches often overemphasize outliers or overly pessimistic scenarios, compromising both overall performance and fairness. To overcome these limitations, we introduce SPECTRE, a minimax-fair method that adjusts the spectrum of a simple Fourier feature mapping and constrains the extent to which the worst-case distribution can deviate from the empirical distribution. We perform extensive experiments on the American Community Survey datasets involving 20 states. The safeness of SPECTRE comes as it provides the highest average values on fairness guarantees together with the smallest interquartile range in comparison to state-of-the-art approaches, even compared to those with access to demographic group information. In addition, we provide a theoretical analysis that derives computable bounds on the worst-case error for both individual groups and the overall population, as well as characterizes the worst-case distributions responsible for these extremal performances
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。