arXiv:2410.08457cs.DCcs.LG2024-10中稿 · TMC, 16 Pages, 12 …被引 2

让弱设备协作训练大模型,提升效率与精度。

Unity is Power: Semi-Asynchronous Collaborative Training of Large-Scale Models with Structured Pruning in Resource-Limited Clients

  • 设计半异步框架,按数据分布剪枝并跨块传知识
  • 精度最高提升8.8%,内存降22%,训练快24%
  • 适合边缘设备、低算力场景的分布式训练

本文研究如何协同利用大量异构且算力弱的设备,在分散数据上训练大规模模型。为在资源受限条件下兼顾效率与精度,首次同时考虑非结构化剪枝、可变子模型架构、知识损失和延迟节点等问题。提出新型半异步协同训练框架 Co-S²P,采用数据分布感知的结构化剪枝与跨块知识迁移机制。理论证明该框架可达到最优收敛率 $O(1/ ext{sqrt}(N^*EQ))$。在真实物联网硬件测试平台上对两类任务进行实验,结果表明:相比现有最佳方法,Co-S²P 在所有资源受限设备上将精度提升最高达8.8%,资源利用率提高1.2倍,内存消耗降低约22%,训练时间减少约24%。

原文摘要 · Abstract (English)

In this work, we study to release the potential of massive heterogeneous weak computing power to collaboratively train large-scale models on dispersed datasets. In order to improve both efficiency and accuracy in resource-adaptive collaborative learning, we take the first step to consider the \textit{unstructured pruning}, \textit{varying submodel architectures}, \textit{knowledge loss}, and \textit{straggler} challenges simultaneously. We propose a novel semi-asynchronous collaborative training framework, namely ${Co\text{-}S}^2{P}$, with data distribution-aware structured pruning and cross-block knowledge transfer mechanism to address the above concerns. Furthermore, we provide theoretical proof that ${Co\text{-}S}^2{P}$ can achieve asymptotic optimal convergence rate of $O(1/\sqrt{N^*EQ})$. Finally, we conduct extensive experiments on two types of tasks with a real-world hardware testbed including diverse IoT devices.The experimental results demonstrate that $Co\text{-}S^2P$ improves accuracy by up to 8.8\% and resource utilization by up to 1.2$\times$ compared to state-of-the-art methods, while reducing memory consumption by approximately 22\% and training time by about 24\% on all resource-limited devices.

联邦学习模型剪枝边缘计算异步训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。