多用户协同推理中,设备自适应学习最优模型分割,显著降低延迟。
CANS: Accelerating Multiuser Collaborative Edge Inference via Cooperative Autodidactic NeuroSurgeon

- 设备共享反馈,动态调整模型分割以适应网络和设备变化
- 在真实硬件上实测,平均推理延迟降低最高达50%
- 适用于资源受限的移动边缘计算场景,适合多设备协同系统
近年来,基于移动边缘计算(MEC)的协同深度神经网络(DNN)推理成为为资源受限移动设备提供智能服务的有前景方案。典型场景是多用户协同边缘推理,各设备独立划分自身DNN模型,并将后端计算任务通过无线网络卸载至共享边缘服务器。然而,由于无线链路波动和设备能力差异等未知且动态变化的系统条件,确定每个设备的最佳DNN分割仍具挑战性。为此,我们提出协作式自教育神经外科医生(CANS),一种通过在线推理过程中共享信息反馈实现设备自适应学习最优分割的协同边缘推理框架。为应对设备异构性并更好利用离线推理经验,我们引入一种新型联邦线性上下文-不确定性下界算法(FedLinUCB-DW),该算法对同类型设备分组,并利用本地离线早退出推理经验进行在线探索的热启动。此外,我们为FedLinUCB-DW提供了理论保障,推导出其后悔上界。我们在模拟环境和硬件原型系统上验证了方法的有效性。实验表明,相比现有先进基线,CANS实现了更低的推理延迟。尤其在两个边缘设备的原型实验中,所提方法较非协同基线平均推理延迟降低高达50%。
原文摘要 · Abstract (English)
Recently, mobile edge computing (MEC)-enabled collaborative deep neural network (DNN) inference has emerged as a promising approach for delivering intelligent services to resource-constrained mobile devices. A representative scenario is multi-user collaborative edge inference, where distinct devices independently partition their DNN models and offload backend computation to a common edge server over wireless networks. However, determining the optimal DNN partition for each device is challenging due to unknown and time-varying system conditions, including fluctuating wireless links and diverse device capabilities. To address this problem, we propose Cooperative Autodidactic NeuroSurgeon (CANS), a collaborative edge inference framework that enables devices to adaptively learn optimal DNN partitions by sharing informative feedback during online inference. To handle the challenge of device heterogeneity and better leverage offline inference experience, we integrate a novel FedLinUCB-DW algorithm that groups devices of the same type and warm-starts online exploration using local offline early-exit inference experience. Furthermore, we provide theoretical guarantees for FedLinUCB-DW by deriving the regret upper bound. We also validate our method on both a simulated environment and a hardware prototype system. Empirical evaluations demonstrate that CANS achieves lower inference latency compared to state-of-the-art baselines. Especially, in prototype experiments on two edge devices, the proposed CANS reduced average inference latency by up to 50% compared to the non-cooperative baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。