arXiv:2503.16583cs.LGcs.AI2025-03中稿 · the ISQED 2025 con…被引 1

用可解释AI加速近似神经网络设计,节能7倍且精度损失仅1-2%。

Explainable AI-Guided Efficient Approximate DNN Generation for Multi-Pod Systolic Arrays

  • 结合硬件模型与可解释AI,自动识别可近似的神经网络层。
  • 在多节点阵列上实现7倍能效提升,精度损失仅1-2%。
  • 适合需要高效部署近似深度学习的芯片设计与系统优化者。

近似深度神经网络(AxDNN)有望提升实际设备中的能效。其中关键因素之一是使用近似乘法器,但其在CPU和GPU上的仿真难以扩展,严重拖慢整体仿真进程,影响高效近似乘法器的筛选。为此,本文提出XAI-Gen方法,利用新兴硬件加速器(如Google TPU v4)的分析模型与可解释人工智能(XAI),精准识别可近似处理的非关键网络层,并快速发现适配各层的近似乘法器。实验表明,XAI-Gen在保持1-2%精度损失的前提下,能将能耗降低最多7倍。通过神经架构搜索(XAI-NAS)案例研究进一步验证:相比现有先进方法,XAI-NAS实现40%更高的能效,执行时间减少达5倍。

原文摘要 · Abstract (English)

Approximate deep neural networks (AxDNNs) are promising for enhancing energy efficiency in real-world devices. One of the key contributors behind this enhanced energy efficiency in AxDNNs is the use of approximate multipliers. Unfortunately, the simulation of approximate multipliers does not usually scale well on CPUs and GPUs. As a consequence, this slows down the overall simulation of AxDNNs aimed at identifying the appropriate approximate multipliers to achieve high energy efficiency with a minimum accuracy loss. To address this problem, we present a novel XAI-Gen methodology, which leverages the analytical model of the emerging hardware accelerator (e.g., Google TPU v4) and explainable artificial intelligence (XAI) to precisely identify the non-critical layers for approximation and quickly discover the appropriate approximate multipliers for AxDNN layers. Our results show that XAI-Gen achieves up to 7x lower energy consumption with only 1-2% accuracy loss. We also showcase the effectiveness of the XAI-Gen approach through a neural architecture search (XAI-NAS) case study. Interestingly, XAI-NAS achieves 40\% higher energy efficiency with up to 5x less execution time when compared to the state-of-the-art NAS methods for generating AxDNNs.

近似计算可解释AI能效优化神经网络设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。