提出隐私保护下统计异质性测量新方法,提升联邦学习准确性。
Towards Robust Federated Analytics via Differentially Private Measurements of Statistical Heterogeneity
- 设计可差分隐私的异质性测量机制,结合根查找优化参数。
- 在分布式场景下,精度优于传统机制与集中式设置。
- 适用于对数据分布差异敏感的联邦学习任务。
统计异质性用于衡量数据集样本分布的偏斜程度。在差分隐私研究中,使用统计异质性高的数据集会导致显著精度损失。在联邦学习场景下,该问题更为突出。本文探讨了三种最具前景的统计异质性测量方法,推导其精度公式,并同时引入差分隐私保护。通过解析机制结合根查找法确定最优隐私参数。实验验证了主要定理及相关假设,并测试了该机制在不同异质性水平下的鲁棒性。在分布式设置中,该解析机制的精度优于所有包含经典机制和/或集中式设置的组合。所有异质性测量方法在使用异质样本时均未出现显著精度损失。
原文摘要 · Abstract (English)
Statistical heterogeneity is a measure of how skewed the samples of a dataset are. It is a common problem in the study of differential privacy that the usage of a statistically heterogeneous dataset results in a significant loss of accuracy. In federated scenarios, statistical heterogeneity is more likely to happen, and so the above problem is even more pressing. We explore the three most promising ways to measure statistical heterogeneity and give formulae for their accuracy, while simultaneously incorporating differential privacy. We find the optimum privacy parameters via an analytic mechanism, which incorporates root finding methods. We validate the main theorems and related hypotheses experimentally, and test the robustness of the analytic mechanism to different heterogeneity levels. The analytic mechanism in a distributed setting delivers superior accuracy to all combinations involving the classic mechanism and/or the centralized setting. All measures of statistical heterogeneity do not lose significant accuracy when a heterogeneous sample is used.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。