通过剪枝与资源协同优化,提升无线边缘网络中大模型推理效率。
The Larger the Merrier? Efficient Large AI Model Inference in Wireless Edge Networks
- 基于剪枝感知的模型分割策略,实现设备与服务器协同推理。
- 剪枝率与系统延迟、能耗间存在可分析的权衡关系,性能更优。
- 适合资源受限的边缘场景,尤其在异构环境下效果显著。
大规模人工智能模型(LAIM)服务需求增长推动推理范式从传统云部署转向边缘部署,以实现低延迟和隐私保护。本文研究一种剪枝感知的边缘-设备联合推理方案,将预训练的LAIM剪枝并分割为设备端与服务器端子模型。理论分析表明,模型输出失真上界由参数失真决定;通过率失真理论推导出参数失真下界,揭示剪枝率与联合推理性能的关系。基于此,构建联合优化剪枝率、传输功率与计算频率的失真最小化问题,考虑系统延迟、能量与资源约束。提出高效算法求解高度非凸问题。大量仿真验证:参数失真可靠反映输出失真;所提联合设计在推理性能、延迟与能耗间平衡优于全设备或全服务器推理方案;分割点在异构且资源受限环境中对性能优化起关键作用。
原文摘要 · Abstract (English)
The growing demand for large artificial intelligence model (LAIM) services is driving a paradigm shift from traditional cloud-based inference to edge-based inference for low-latency, privacy-preserving applications. In particular, edge-device co-inference, which partitions LAIMs between edge devices and servers, has emerged as a promising strategy for resource-efficient LAIM execution in wireless networks. In this paper, we investigate a pruning-aware LAIM co-inference scheme, where a pre-trained LAIM is pruned and partitioned into on-device and on-server sub-models for deployment. For analysis, we first prove that the LAIM output distortion is upper bounded by its parameter distortion. Then, we derive a lower bound on parameter distortion via rate-distortion theory, analytically capturing the relationship between pruning ratio and co-inference performance. Next, based on the analytical results, we formulate an LAIM co-inference distortion bound minimization problem by jointly optimizing the pruning ratio, transmit power, and computation frequency under system latency, energy, and available resource constraints. Moreover, we propose an efficient algorithm to tackle the considered highly non-convex problem. Finally, extensive simulations demonstrate the effectiveness of the proposed design. In particular, model parameter distortion is shown to provide a reliable bound on output distortion. Also, the proposed joint pruning ratio and resource management design achieves superior performance in balancing trade-offs among inference performance, system latency, and energy consumption compared with benchmark schemes, such as fully on-device and on-server inference. Moreover, the split point is shown to play a critical role in system performance optimization under heterogeneous and resource-limited edge environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。