arXiv:2605.27475cs.LGcs.AI2026-05被引 4

提出HEAL框架,让分布式学习既快速又抗故障

HEAL: Resilient and Self-* Hub-based Learning

  • 用动态选主节点方式融合联邦与泛洪学习优点
  • 无故障时性能媲美联邦学习,故障下收敛更快
  • 适合高故障率场景,如物联网或边缘计算

去中心化学习通过在节点间分布数据和计算提升隐私性、可扩展性和容错能力。主流的联邦学习依赖中心聚合器,面临服务器漏洞、扩展性差、隐私风险及单点故障问题。而泛洪学习和流行病学习虽实现完全去中心化,通过节点间直接交换模型更新保证鲁棒性与隐私,但收敛速度较慢。本文提出新型去中心化学习框架HEAL,是首个跨层设计的框架,结合优化自组织自愈型P2P网络,融合联邦学习、泛洪学习与流行病学习的优势。基于近期提出的电梯算法(Elevator algorithm),HEAL动态选择节点担任聚合角色。仿真表明,在无故障环境下,HEAL性能接近联邦学习;在存在崩溃与节点频繁变动的环境中,其表现优于泛洪学习和流行病学习。

原文摘要 · Abstract (English)

Decentralized learning enhances privacy, scalability, and fault tolerance by distributing data and computation across nodes. A popular approach is Federated learning, which relies on a central aggregator, yet faces challenges such as server vulnerabilities, scalability issues, privacy risks and most importantly, the single point of failure. Alternatively Gossip Learning and Epidemic Learning offer fully decentralization through peer-to-peer exchanges of model updates, ensuring robustness and privacy, at the price of slower model convergence. In this work, we introduce a novel decentralized learning framework called HEAL. HEAL is the first cross-layer decentralized learning framework that exploits an optimized self-organizing and self-healing underlying P2P overlay combining the strengths of Federated Learning, Gossip and Epidemic Learning. Leveraging the recently proposed Elevator algorithm, HEAL promotes dynamically chosen nodes to act as aggregators. Through simulations, we demonstrate that HEAL has similar performances to that of Federated Learning in crash-free settings, while being fully decentralized and fault-tolerant. In crash and churn prone environments HEAL outperforms Gossip and Epidemic Learning.

去中心化学习联邦学习容错系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。