跨HPC与云环境的联邦学习框架,解决异构硬件下的隐私计算难题
Federated Learning Framework for Scalable AI in Heterogeneous HPC and Cloud Environments
- 设计跨HPC与云的联邦学习框架,支持异构资源协同训练
- 在非独立同分布数据下仍保持收敛性与高准确率
- 适合需要隐私保护的超大规模分布式AI系统研发人员
随着对可扩展、隐私敏感的AI系统需求增长,联邦学习(FL)成为一种有前景的解决方案,可在不移动原始数据的前提下实现去中心化模型训练。与此同时,高性能计算(HPC)与云基础设施的结合提供了巨大算力,但也带来了异构硬件、通信瓶颈和非均匀数据等新挑战。本文提出一个专为混合HPC与云环境设计的联邦学习框架,有效应对系统异构性、通信开销和资源调度问题,同时保障模型精度与数据隐私。在混合测试平台上实验表明,该系统在可扩展性、容错性和收敛性方面表现优异,即使在非独立同分布(non-IID)数据和多样化硬件条件下依然稳定运行。结果表明,联邦学习是构建现代分布式计算环境中可扩展AI系统的实用路径。
原文摘要 · Abstract (English)
As the demand grows for scalable and privacy-aware AI systems, Federated Learning (FL) has emerged as a promising solution, allowing decentralized model training without moving raw data. At the same time, the combination of high-performance computing (HPC) and cloud infrastructure offers vast computing power but introduces new complexities, especially when dealing with heterogeneous hardware, communication limits, and non-uniform data. In this work, we present a federated learning framework built to run efficiently across mixed HPC and cloud environments. Our system addresses key challenges such as system heterogeneity, communication overhead, and resource scheduling, while maintaining model accuracy and data privacy. Through experiments on a hybrid testbed, we demonstrate strong performance in terms of scalability, fault tolerance, and convergence, even under non-Independent and Identically Distributed (non-IID) data distributions and varied hardware. These results highlight the potential of federated learning as a practical approach to building scalable Artificial Intelligence (AI) systems in modern, distributed computing settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。