提出稳定高效的联邦学习优化器,解决非独立同分布数据下的训练不稳问题。
Taming the Instability: A Robust Second-Order Optimizer for Federated Learning over Non-IID Data
- 引入梯度异常监测与故障安全重置机制,实时应对数值不稳定性。
- 在多种非独立同分布场景下,收敛速度更快且准确率更高。
- 适合需要高鲁棒性的分布式机器学习应用,如医疗、金融等隐私敏感领域。
本文提出联邦鲁棒曲率优化(FedRCO),一种新型二阶优化框架,旨在改善统计异质性下联邦学习的收敛速度并降低通信开销。现有二阶优化方法在分布式环境中常因计算成本高和数值不稳定性而受限。相比之下,FedRCO通过集成高效近似曲率优化器与可证明稳定的机制来克服这些挑战。具体包含三个关键组件:(1) 梯度异常监测器,实时检测并缓解梯度爆炸;(2) 故障安全韧性协议,在出现数值不稳定性时重置优化状态;(3) 曲率保持自适应聚合策略,安全融合全局知识而不破坏局部曲率结构。理论分析表明,FedRCO能有效抑制不稳定性,防止无界更新,同时保持优化效率。大量实验显示,该方法在多种非独立同分布场景中表现出更强鲁棒性,相比最先进的首阶与二阶方法,实现了更高的准确率和更快的收敛速度。
原文摘要 · Abstract (English)
In this paper, we present Federated Robust Curvature Optimization (FedRCO), a novel second-order optimization framework designed to improve convergence speed and reduce communication cost in Federated Learning systems under statistical heterogeneity. Existing second-order optimization methods are often computationally expensive and numerically unstable in distributed settings. In contrast, FedRCO addresses these challenges by integrating an efficient approximate curvature optimizer with a provable stability mechanism. Specifically, FedRCO incorporates three key components: (1) a Gradient Anomaly Monitor that detects and mitigates exploding gradients in real-time, (2) a Fail-Safe Resilience protocol that resets optimization states upon numerical instability, and (3) a Curvature-Preserving Adaptive Aggregation strategy that safely integrates global knowledge without erasing the local curvature geometry. Theoretical analysis shows that FedRCO can effectively mitigate instability and prevent unbounded updates while preserving optimization efficiency. Extensive experiments show that FedRCO achieves superior robustness against diverse non-IID scenarios while achieving higher accuracy and faster convergence than both state-of-the-art first-order and second-order methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。