按区域暴露程度动态分配隐私噪声,提升联邦学习精度与隐私平衡。
Topology-Aware Differential Privacy in Hierarchical Federated Learning
- 基于区域结构设计差异隐私噪声分配策略,避免过度加噪。
- 在ε=0.99下,图像与文本任务准确率最高提升14.84%和12.16%。
- 适用于异构部署的联邦学习系统,尤其适合区域暴露不均场景。
层级联邦学习在客户端与云端之间设置区域聚合器,使得每个参与者的更新仅与其邻居可见。这种隐藏机制的效果取决于聚合区域的大小,而实际部署中各区域规模差异显著。当前做法对所有参与者统一添加噪声,按最暴露区域校准,导致其他参与者承受超出其实际暴露需求的噪声。本文提出显式解决方案:首先给出分层级差分隐私保证,再通过匹配保护目标的邻接关系,界定了参与者本地类别分布与上层观察者估计之间的互信息。在固定效用预算下最小化最坏情况边界,得到一种名为Fulcrum的极小极大最优分配方案。该方案恢复的预算具有闭式表达,称为暴露离散度,衡量区域内聚合权重相对于最暴露区域的集中程度。该值仅由区域结构和聚合权重决定,可在训练前评估,当所有区域暴露程度相等时精确为零。在图像与文本分类任务中,ε=0.99时,暴露离散度大时,准确率分别提升最多14.84和12.16个百分点;在均衡控制组中,提升恰好为零,与理论预测一致。
原文摘要 · Abstract (English)
Hierarchical federated learning places regional aggregators between clients and the cloud, so a participant's update is observed only alongside its neighbours'. The concealment this arrangement provides depends on the size of the aggregation region, and regions in operational deployments vary widely. Prevailing practice applies a single noise multiplier to every participant, calibrated for the most exposed region, so every other participant carries more noise than its own exposure requires. We show that this allocation problem admits an explicit solution. We first give a silo-level differential privacy guarantee for the mechanism, then bound the mutual information between a participant's local class distribution and any estimate an observer positioned above the regional tier could form of it, using an adjacency notion matched to the quantity being protected. Minimising the worst-case bound under a fixed utility budget yields a min-max optimal allocation, which we call Fulcrum. The budget it recovers has a closed form we term the exposure dispersion, a measure of how unevenly aggregation weight is concentrated within regions relative to the most exposed one. Because this quantity follows from the region structure and the aggregation weights alone, a practitioner can evaluate it before training begins, and it vanishes precisely when all regions are equally exposed. On image and text classification at $\varepsilon = 0.99$, accuracy at a matched worst-case per-client guarantee improves by up to $14.84$ and $12.16$ percentage points where the dispersion is large, and is exactly zero on a balanced control for which the theory predicts parity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。