在联邦学习中同时保障隐私与公平性,提出高效算法解决数据分布下的偏见问题。
Federated fairness-aware classification under differential privacy
- 分两步设计联邦隐私公平分类算法,分离隐私与公平的代价来源。
- 理论证明算法在隐私、公平和风险控制上均有保障,误差可控。
- 适用于多服务器场景,兼顾计算效率与实际应用价值。
隐私与算法公平性已成为现代机器学习的两大核心议题。尽管两者各自发展迅速,但其联合影响仍研究不足。本文系统研究联邦设置下差分隐私与公平性对分类的影响,针对受人口统计差异约束的联邦差分隐私分类问题,提出两步算法FDP-Fair;在单服务器情况下,进一步提出轻量级算法CDP-Fair。在合理结构假设下,建立了隐私、公平与过拟合风险的理论保证。特别地,将私有公平分类的额外风险分解为:(a) 分类固有成本,(b) 私有分类成本,(c) 非私有公平成本,(d) 私有公平成本。理论结果通过合成与真实数据集上的大量实验得到验证,凸显所提算法的实用性。
原文摘要 · Abstract (English)
Privacy and algorithmic fairness have become two central issues in modern machine learning. Although each has separately emerged as a rapidly growing research area, their joint effect remains comparatively under-explored. In this paper, we systematically study the joint impact of differential privacy and fairness on classification in a federated setting, where data are distributed across multiple servers. Targeting demographic disparity constrained classification under federated differential privacy, we propose a two-step algorithm, namely FDP-Fair. In the special case where there is only one server, we further propose a simple yet powerful algorithm, namely CDP-Fair, serving as a computationally-lightweight alternative. Under mild structural assumptions, theoretical guarantees on privacy, fairness and excess risk control are established. In particular, we disentangle the source of the private fairness-aware excess risk into a) intrinsic cost of classification, b) cost of private classification, c) non-private cost of fairness and d) private cost of fairness. Our theoretical findings are complemented by extensive numerical experiments on both synthetic and real datasets, highlighting the practicality of our designed algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。