arXiv:2506.21371cs.LGcs.AR2025-06被引 7

通过多层次近似计算,实现深度神经网络硬件的低功耗高效推理。

MAx-DNN: Multi-Level Arithmetic Approximation for Energy-Efficient DNN Hardware Accelerators

  • 在层、滤波器和核级别精细部署近似乘法器,平衡精度与能耗。
  • 相比基线量化模型,能效提升54%,精度损失不超过4%。
  • 相较现有方法,能效翻倍且精度更高,适合边缘设备部署。

当前深度神经网络(DNN)架构的快速发展使其成为提供先进机器学习任务的主流方案,具备优异的准确率。为实现低功耗DNN计算,本文研究了DNN工作负载的细粒度误差容错性与硬件近似技术的协同作用,以提升能效。基于最新的ROUP近似乘法器,我们采用层级、滤波器级和核级的策略系统地探索其在网络中的分布,并评估其对准确率和能耗的影响。在CIFAR-10数据集上的ResNet-8模型实验表明,所提方案相较于基线量化模型,最多可实现54%的能效提升,精度损失不超过4%;同时相比现有最先进近似方法,能效提升2倍且准确率更优。

原文摘要 · Abstract (English)

Nowadays, the rapid growth of Deep Neural Network (DNN) architectures has established them as the defacto approach for providing advanced Machine Learning tasks with excellent accuracy. Targeting low-power DNN computing, this paper examines the interplay of fine-grained error resilience of DNN workloads in collaboration with hardware approximation techniques, to achieve higher levels of energy efficiency. Utilizing the state-of-the-art ROUP approximate multipliers, we systematically explore their fine-grained distribution across the network according to our layer-, filter-, and kernel-level approaches, and examine their impact on accuracy and energy. We use the ResNet-8 model on the CIFAR-10 dataset to evaluate our approximations. The proposed solution delivers up to 54% energy gains in exchange for up to 4% accuracy loss, compared to the baseline quantized model, while it provides 2x energy gains with better accuracy versus the state-of-the-art DNN approximations.

神经网络加速低功耗设计近似计算硬件优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。