用几何方法重新定义联邦学习聚合,提升模型一致性与可靠性。
Information-Geometric Barycenters for Bayesian Federated Learning
- 将联邦学习聚合视为后验分布的几何中位点求解
- 在非独立同分布场景下性能媲美顶尖方法
- 适合关注模型不确定性与理论解释的研究者
联邦学习(FL)是一种广泛使用的分布式优化框架,通过平均本地训练模型实现共识。然而,该方法与贝叶斯推断的结构不一致,后者要求模型空间具有分布结构。本文从信息几何角度重新诠释联邦学习聚合问题:将其视为在预设散度度量下,最小化各客户端间平均偏差的局部后验分布的几何中位点求解。这一视角为众多现有方法提供了统一框架,并揭示其理论基础。我们提出BA-BFL算法,在非凸环境下保持联邦平均的收敛性。在非独立同分布设置下,与多种统计聚合技术进行广泛对比,结果显示BA-BFL性能接近当前最优方法,同时提供聚合过程的几何解释。此外,我们还将分析扩展至混合贝叶斯深度学习,探究贝叶斯层对不确定性量化与模型校准的影响。
原文摘要 · Abstract (English)
Federated learning (FL) is a widely used and impactful distributed optimization framework that achieves consensus through averaging locally trained models. While effective, this approach may not align well with Bayesian inference, where the model space has the structure of a distribution space. Taking an information-geometric perspective, we reinterpret FL aggregation as the problem of finding the barycenter of local posteriors using a prespecified divergence metric, minimizing the average discrepancy across clients. This perspective provides a unifying framework that generalizes many existing methods and offers crisp insights into their theoretical underpinnings. We then propose BA-BFL, an algorithm that retains the convergence properties of Federated Averaging in non-convex settings. In non-independent and identically distributed scenarios, we conduct extensive comparisons with statistical aggregation techniques, showing that BA-BFL achieves performance comparable to state-of-the-art methods while offering a geometric interpretation of the aggregation phase. Additionally, we extend our analysis to Hybrid Bayesian Deep Learning, exploring the impact of Bayesian layers on uncertainty quantification and model calibration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。