arXiv:2509.17970cs.LGcs.AI2025-09被引 4

同时调节内存与计算频率,显著降低边缘设备的DNN推理能耗。

Joint Memory Frequency and Computing Frequency Scaling for Energy-efficient DNN Inference

  • 联合调整内存与计算频率,实现能效优化。
  • 仿真显示协同调频可降低设备能耗,适用于本地与协同推理场景。
  • 为资源受限设备提供低延迟、低功耗的DNN推理方案,适合边缘计算开发者。

深度神经网络(DNN)已在多种应用中广泛部署,但在资源受限设备上仍面临高延迟和高能耗问题。现有研究多聚焦于通过动态电压频率调节(DVFS)调整处理器计算频率以平衡延迟与能耗,但内存频率的调整常被忽视,未能充分发挥其在推理时间和能耗中的作用。本文首次采用基于模型与数据驱动的方法,研究联合调节内存频率与计算频率对推理时间与能耗的影响。结合不同DNN模型的拟合参数,初步分析了同步调整内存与计算频率的效果。仿真结果在本地推理与协同推理场景下均验证了联合调频的有效性,能够显著降低设备能耗。

原文摘要 · Abstract (English)

Deep neural networks (DNNs) have been widely applied in diverse applications, but the problems of high latency and energy overhead are inevitable on resource-constrained devices. To address this challenge, most researchers focus on the dynamic voltage and frequency scaling (DVFS) technique to balance the latency and energy consumption by changing the computing frequency of processors. However, the adjustment of memory frequency is usually ignored and not fully utilized to achieve efficient DNN inference, which also plays a significant role in the inference time and energy consumption. In this paper, we first investigate the impact of joint memory frequency and computing frequency scaling on the inference time and energy consumption with a model-based and data-driven method. Then by combining with the fitting parameters of different DNN models, we give a preliminary analysis for the proposed model to see the effects of adjusting memory frequency and computing frequency simultaneously. Finally, simulation results in local inference and cooperative inference cases further validate the effectiveness of jointly scaling the memory frequency and computing frequency to reduce the energy consumption of devices.

DNN推理能效优化边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。