arXiv:2411.16086cs.DCcs.AI2024-11中稿 · be published in 28…被引 16

针对异构边缘设备,提出分层分割策略降低推理延迟。

HiDP: Hierarchical DNN Partitioning for Distributed Inference on Heterogeneous Edge Platforms

  • 分层设计:全局与局部双重分割,适配设备核心差异
  • 平均延迟降38%,能耗降46%,吞吐量提56%
  • 适合部署于多核异构边缘设备的智能推理系统

边缘推理技术将深度神经网络(DNN)任务在多个边缘节点间划分与分发,以实现低延迟推理,但未考虑边缘节点的核级异构性。此外,默认的DNN推理框架也未能充分利用异构边缘节点的资源,导致推理延迟更高。本文提出一种面向异构边缘节点的分层DNN分割策略(HiDP)。该策略通过考虑边缘节点的核级异构性,在全局和局部层面进行分层分割DNN工作负载。我们在商用边缘设备上,对多种主流DNN模型进行了评估。结果表明,与现有相关方法相比,所提策略平均降低了38%的延迟、46%的能耗,并提升了56%的吞吐量。

原文摘要 · Abstract (English)

Edge inference techniques partition and distribute Deep Neural Network (DNN) inference tasks among multiple edge nodes for low latency inference, without considering the core-level heterogeneity of edge nodes. Further, default DNN inference frameworks also do not fully utilize the resources of heterogeneous edge nodes, resulting in higher inference latency. In this work, we propose a hierarchical DNN partitioning strategy (HiDP) for distributed inference on heterogeneous edge nodes. Our strategy hierarchically partitions DNN workloads at both global and local levels by considering the core-level heterogeneity of edge nodes. We evaluated our proposed HiDP strategy against relevant distributed inference techniques over widely used DNN models on commercial edge devices. On average our strategy achieved 38% lower latency, 46% lower energy, and 56% higher throughput in comparison with other relevant approaches.

边缘计算异构计算推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。