跨超算中心联邦学习框架,实现科学大模型协同训练
Scalable Cross-Facility Federated Learning for Scientific Foundation Models on Multiple Supercomputers
- 基于APPFL与Globus构建跨设施联邦学习框架
- 在4台美国能源部超算上验证可行,算法选择影响显著
- 适合需隐私保护的科学领域大模型协作研究
科学领域的人工智能应用日益依赖大规模模型训练,但数据因隐私、主权或体量限制无法集中。联邦学习(FL)可在不集中原始数据的前提下实现协同训练,但科学模型规模巨大,需依托高性能计算(HPC)设施。跨HPC设施部署FL面临超越云环境的挑战。本文提出一个面向异构HPC环境的完整跨设施联邦学习框架,基于先进隐私保护联邦学习(APPFL)架构,结合Globus Compute与传输编排技术,并在四台美国能源部领导级超算上评估。结果表明,跨超算的联邦学习实验在实际中可实现,揭示了影响训练性能的关键异质性来源,并证明在真实HPC调度条件下算法选择至关重要。通过在化学指令数据集上微调大型语言模型验证了其科学适用性,指出调度感知的算法设计是未来部署的关键开放挑战。
原文摘要 · Abstract (English)
Artificial Intelligence for scientific applications increasingly requires training large models on data that cannot be centralized due to privacy constraints, data sovereignty, or the sheer volume of data generated. Federated learning (FL) addresses this by enabling collaborative training without centralizing raw data, but scientific applications demand model scales that requires extensive computing resources, typically offered at High Performance Computing (HPC) facilities. Deploying FL experiments across HPC facilities introduces challenges beyond cloud or enterprise settings. We present a comprehensive cross-facility FL framework for heterogeneous HPC environments, built on Advanced Privacy-Preserving Federated Learning (APPFL) framework with Globus Compute and Transfer orchestration, and evaluate it across four U.S. Department of Energy (DOE) leadership-class supercomputers. We demonstrate that FL experiments across HPC facilities are practically achievable, characterize key sources of heterogeneity impacting the training performance, and show that algorithmic choices matter significantly under realistic HPC scheduling conditions. We validate the scientific applicability by fine-tuning a large language model on a chemistry instruction dataset, and identify scheduler-aware algorithm design as a critical open challenge for future deployments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。