arXiv:2608.13863cs.AI2026-08

同时优化计算与内存频率,让手机端模型推理更省电

Joint Optimization of Memory and Computing Frequency for Energy-Efficient DNN Inference

  • 联合调节计算频率和内存频率以降低能耗
  • 本地推理可实现接近最优的能效,误差仅2.5%以内
  • 适合资源受限的移动端模型部署场景

移动设备上的深度神经网络(DNN)推理常因计算和内存资源有限而产生高延迟和高功耗。现有研究多关注通过动态电压频率调节(DVFS)调整计算频率,却忽视了内存频率的影响。本文考虑计算频率与内存频率对推理时延的共同影响,联合优化两者及通信资源以实现能效最优。基于真实推理时延模型,构建了在截止时间约束下最小化所有设备能耗的优化问题。针对本地推理,通过凸优化推导出近似最优闭式解;针对边缘推理,在给定带宽下获得传输功率的最优闭式解。此外,提出一种低复杂度启发式算法,具有多项式时间复杂度。基于实测数据的仿真结果表明,本地推理的近优解在严格截止时间内表现接近最优,性能差距不超过2.5%;所提算法相比其他方法最多可降低10.4%的设备能耗。

原文摘要 · Abstract (English)

Deep neural network (DNN) inference on mobile devices often incurs high latency and energy consumption due to limited computing and memory resources. To enable energy-efficient DNN inference, most existing studies focus on dynamic voltage and frequency scaling (DVFS) for adjusting the computing frequency, while the impact of memory frequency on the inference performance has been greatly overlooked. In this paper, we consider the impact of memory frequency and computing frequency on DNN inference time, and jointly optimize these two frequencies together with communication resources for energy-efficient DNN inference. Based on a realistic inference time model, we formulate an optimization problem to minimize the energy consumption of all mobile devices under the deadline constraint. For local inference, we derive a near-optimal closed-form solution via convex optimization, while an optimal closed-form solution for transmission power is obtained for edge inference with the given bandwidth. Furthermore, we propose a low-complexity heuristic algorithm to effectively solve the overall problem with polynomial time complexity. Simulation results based on measured data show that the proposed near-optimal solution for local inference can achieve optimal performance under strict deadline constraints, with a performance gap of up to 2.5% compared with the optimal solution. Meanwhile, our proposed algorithm significantly reduces the energy consumption of devices by up to 10.4% compared to other methods.

能效优化DNN推理移动端频率调节

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。