用分位数摘要实现跨数据孤岛的公平性审计,不共享原始数据。
Federated Measurement of Demographic Disparities from Quantile Sketches
- 通过分位数摘要和瓦瑟斯坦距离,实现隐私保护下的群体公平性测量。
- 仅需几十个分位点,即可准确估算全局公平性差异及来源。
- 适合关注模型公平性但受限于数据隐私的研究者与开发者。
许多公平性目标在总体层面定义,与受隐私法规限制而无法共享的分布式数据收集不匹配。横向联邦学习(FL)可在不共享原始数据的前提下,实现特征对齐的客户端协作建模。本文研究基于评分分布的群体公平性审计,将差异度量为敏感群体评分分布间的瓦瑟斯坦-弗雷歇方差,并将总体指标表达为联邦形式,明确揭示了局部选择机制如何导致本地与全局结果的偏差。针对平方瓦瑟斯坦距离,我们证明了类似ANOVA的分解,可分离出(i)由选择引入的混合效应与(ii)跨客户端异质性,从而得到紧致的局部-全局指标关联边界。随后提出一种单轮、通信高效的协议:每个数据孤岛仅需共享组别数量及其本地评分分布的分位数摘要,服务器即可估计全局差异及其分解,具有O(1/k)的离散化偏差(k个分位点)和有限样本保证。在合成数据与COMPAS数据集上的实验表明,仅需几十个分位点即可准确恢复全局差异并诊断其来源。
原文摘要 · Abstract (English)
Many fairness goals are defined at a population level that misaligns with siloed data collection, which remains unsharable due to privacy regulations. Horizontal federated learning (FL) enables collaborative modeling across clients with aligned features without sharing raw data. We study federated auditing of demographic parity through score distributions, measuring disparity as a Wasserstein--Frechet variance between sensitive-group score laws, and expressing the population metric in federated form that makes explicit how silo-specific selection drives local-global mismatch. For the squared Wasserstein distance, we prove an ANOVA-style decomposition that separates (i) selection-induced mixture effects from (ii) cross-silo heterogeneity, yielding tight bounds linking local and global metrics. We then propose a one-shot, communication-efficient protocol in which each silo shares only group counts and a quantile summary of its local score distributions, enabling the server to estimate global disparity and its decomposition, with $O(1/k)$ discretization bias ($k$ quantiles) and finite-sample guarantees. Experiments on synthetic data and COMPAS show that a few dozen quantiles suffice to recover global disparity and diagnose its sources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。