arXiv:2601.08135cs.NIcs.DC2026-01

提出分层调度框架,实现边缘推理中能效与精度的协同优化。

Hierarchical Online-Scheduling for Energy-Efficient Split Inference with Progressive Transmission

  • 分两层调度:任务级动态划分模型,包级自适应传输应对信道波动。
  • 严苛延迟下精度提升43.12%,能耗降低62.13%,多用户场景稳定高效。
  • 适合需要低功耗高精度边缘推理的物联网与移动设备应用。

基于深度神经网络(DNN)的设备-边缘协同推理面临精度、延迟与能耗之间的根本权衡。现有调度存在两大缺陷:粗粒度任务级决策与细粒度包级信道动态之间存在粒度不匹配,且对每项任务的复杂度认知不足,导致仅在任务级调度时资源利用效率低下。本文提出新型能量-精度分层优化框架ENACHI,联合优化任务级与包级调度,在满足能耗与延迟约束下最大化推理精度。设计双层李雅普诺夫框架,引入渐进式传输技术以增强适应性。任务级通过外层漂移-惩罚环在线决策DNN分割与带宽分配,并建立参考功耗预算以管理长期能效-精度权衡;包级采用不确定性感知的渐进式传输机制,动态应对样本级任务复杂度变化,并结合嵌套内控制环实施新型参考跟踪策略,实时调整每时隙发射功率以适应信道波动。ImageNet实验表明,相较于最先进基准,ENACHI在不同截止时间与带宽条件下表现更优,严苛延迟下精度提升43.12%,能耗降低62.13%,且在多用户拥塞场景中保持稳定的能耗表现,展现出优异可扩展性。

原文摘要 · Abstract (English)

Device-edge collaborative inference with Deep Neural Networks (DNNs) faces fundamental trade-offs among accuracy, latency and energy consumption. Current scheduling exhibits two drawbacks: a granularity mismatch between coarse, task-level decisions and fine-grained, packet-level channel dynamics, and insufficient awareness of per-task complexity. Consequently, scheduling solely at the task level leads to inefficient resource utilization. This paper proposes a novel ENergy-ACcuracy Hierarchical optimization framework for split Inference, named ENACHI, that jointly optimizes task- and packet-level scheduling to maximize accuracy under energy and delay constraints. A two-tier Lyapunov-based framework is developed for ENACHI, with a progressive transmission technique further integrated to enhance adaptivity. At the task level, an outer drift-plus-penalty loop makes online decisions for DNN partitioning and bandwidth allocation, and establishes a reference power budget to manage the long-term energy-accuracy trade-off. At the packet level, an uncertainty-aware progressive transmission mechanism is employed to adaptively manage per-sample task complexity. This is integrated with a nested inner control loop implementing a novel reference-tracking policy, which dynamically adjusts per-slot transmit power to adapt to fluctuating channel conditions. Experiments on ImageNet dataset demonstrate that ENACHI outperforms state-of-the-art benchmarks under varying deadlines and bandwidths, achieving a 43.12\% gain in inference accuracy with a 62.13\% reduction in energy consumption under stringent deadlines, and exhibits high scalability by maintaining stable energy consumption in congested multi-user scenarios.

边缘计算模型分割能效优化动态调度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。