arXiv:2411.16478cs.LGcs.DB2024-11被引 1

在隐私保护下高效估算联邦数据分布差异

Federated and differentially private estimation of KL divergence

  • 基于差分隐私设计新算法,避免直接共享数据
  • 实验显示精度接近非私有方法,且更稳定
  • 适合需要隐私保护的联邦学习与数据分析场景

衡量分布漂移是管理分布式敏感数据的关键任务,支撑了广泛的联邦学习与分析应用。然而,在实际场景中,直接共享此类信息往往不可取(如隐私顾虑)或不可行(如通信成本高)。本文提出 FedPriKL,一种在差分隐私(DP)保障下估计联邦模型间数据KL散度的新方法。我们建立了其理论性质,证明该方法无偏且敏感度与方差低且有界,从而在DP条件下仍保持强实用性。此外,我们通过实证研究探索参数选择以优化精度并最小化通信开销。实验表明,FedPriKL的精度与类似的采样非私有估计器相当,且在稳定性与准确性上优于用户本地扰动数据的基线方法。

原文摘要 · Abstract (English)

Measuring distribution drifts is a key task in managing distributed, sensitive data, as it underpins a wide range of federated learning and analytics applications. In many practical settings, however, directly sharing such information is either undesirable (e.g., due to privacy concerns) or infeasible (e.g., due to high communication costs). In this work, we present FedPriKL, a novel method for estimating the KL divergence of data across federated computational models under differential privacy (DP) guarantees. We establish its theoretical properties, showing that FedPriKL is unbiased with low and bounded sensitivity and variance, thereby ensuring strong utility under DP. In addition, we present an empirical study that explores parameter choices to optimize accuracy while minimizing communication overhead. Our experiments demonstrate that FedPriKL achieves accuracy comparable to a similar sampling-based non-private estimator, while delivering greater stability and accuracy than a baseline variant in which users perturb their data locally to obtain privacy guarantees.

联邦学习差分隐私KL散度分布检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。