arXiv:2605.03499cs.LGcs.IT2026-05

提出分层采样框架,用Wasserstein距离更精确地界定联邦学习泛化误差。

A Hierarchical Sampling Framework for bounding the Generalization Error of Federated Learning

论文配图:A Hierarchical Sampling Framework for bounding the Generalization Error of Federated Learning
图 1 · 摘自论文原文
  • 构建多层树结构建模客户端数据依赖,分层采样设计更贴近真实场景。
  • 在损失函数Lipschitz条件下,导出基于Wasserstein距离的泛化界,理论更紧致。
  • 适用于隐私保护分析,可与差分隐私结合,适合关注安全与泛化权衡的研究者。

我们研究基于Wasserstein距离的分层联邦学习(HFL)预期泛化界。提出一种广义框架,其中数据以分层方式采样,并用多层树结构建模客户端数据间的依赖关系。在损失函数满足Lipschitz假设下,通过超样本构造方法,量化算法对采样树中单个节点变化的敏感性,推导出基于Wasserstein距离的泛化界。利用联邦学习结构,我们恢复并严格优于现有最优条件互信息(CMI)界(在有界损失情形)。此外,证明该界可与差分隐私假设结合,导出基于算法隐私的泛化界。通过高斯位置模型(GLM)验证,我们的界能恢复泛化误差的真实渐近率。

原文摘要 · Abstract (English)

We study expected generalization bounds for the Hierarchical Federated Learning (HFL) setup using Wasserstein distance. We introduce a generalized framework in which data is sampled hierarchically, and we model it with a multi-layered tree structure that induces dependencies among the clients' datasets. We derive generalization bounds in terms of Wasserstein distance under the Lipschitz assumption on the loss function, by applying a supersample construction that allows us to measure the sensitivity of the algorithm to the change of a single node in the sampling tree. By leveraging the FL structure, we recover and strictly imply existing state-of-the-art conditional mutual information (CMI) bounds in the case of bounded losses. We also show that our bound can be applied together with Differential Privacy assumptions, to recover generalization bounds based on algorithmic privacy. To assess the tightness of our bounds, we study the Gaussian Location Model (GLM) and show that we recover the actual asymptotic rate of the generalization error.

联邦学习泛化误差Wasserstein隐私分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。