arXiv:2504.09114cs.LG2025-04被引 3

让大模型在手机等设备上高效运行,兼顾隐私与性能。

Deploying Large AI Models on Resource-Limited Devices with Split Federated Learning

  • 将模型训练拆分到设备和服务器,降低设备内存负担。
  • 通过量化和资源调度,提升训练效率并减少能耗与延迟。
  • 适合移动端、边缘计算等资源受限场景的AI部署。

大规模人工智能模型(LAMs)凭借海量数据、庞大参数和高算力需求,在多个领域带来变革。然而,其在资源受限的移动边缘设备上的实际部署面临数据隐私、资源约束和高开销等挑战。为此,本文提出一种新框架——量化分层联邦微调大模型(SFLAM)。该框架采用分层学习范式,将训练负载在边缘设备与服务器间分割,使大型模型可在设备端运行,并显著降低边缘设备的内存需求。同时,SFLAM引入量化管理、功率控制与带宽分配策略,提升训练效率,同步减少能量消耗与通信延迟。论文还进行了关于时延-能耗权衡的理论分析,并通过全面仿真验证了其有效性。结果表明,相比传统方法,SFLAM在学习效率与可扩展性方面表现更优,为资源受限场景下的先进AI服务提供了可行路径。

原文摘要 · Abstract (English)

Large Artificial Intelligence Models (LAMs) powered by massive datasets, extensive parameter scales, and extensive computational resources, leading to significant transformations across various industries. Yet, their practical deployment on resource-limited mobile edge devices is hindered by critical challenges such as data privacy, constrained resources, and high overhead costs. Addressing this gap, this paper proposes a novel framework, named Quantized Split Federated Fine-Tuning Large AI Model (SFLAM). By partitioning the training load between edge devices and servers using a split learning paradigm, SFLAM can facilitate the operation of large models on devices and significantly lowers the memory requirements on edge devices. Additionally, SFLAM incorporates quantization management, power control, and bandwidth allocation strategies to enhance training efficiency while concurrently reducing energy consumption and communication latency. A theoretical analysis exploring the latency-energy trade-off is presented, and the framework's efficacy is validated via comprehensive simulations. The findings indicate that SFLAM achieves superior performance in terms of learning efficiency and scalability compared to conventional methods, thereby providing a valuable approach for enabling advanced AI services in resource-constrained scenarios.

边缘计算大模型部署联邦学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。