arXiv:2410.23082cs.ARcs.AI2024-10被引 2

提出可灵活配置精度的存内计算芯片,显著降低神经网络推理能耗。

An Event-Based Digital Compute-In-Memory Accelerator with Flexible Operand Resolution and Layer-Wise Weight/Output Stationarity

  • 支持任意数据精度和形状,统一存储权重与膜电位
  • 实现混合权重与输出静态数据流,能效提升2倍
  • 适合边缘视觉场景,节能高达90%且准确率达95.8%

面向脉冲神经网络(SNNs)的存内计算(CIM)加速器有望在边缘视觉应用中实现微秒级延迟与超低功耗。然而,当前电路与系统层面灵活性不足,限制了其在真实场景中的部署。本文提出一种新型数字CIM宏,支持任意操作数分辨率与形状,并采用统一存储单元同时存放权重与膜电位。该电路设计使系统层实现混合权重-输出静态数据流,最大化操作数复用,显著减少芯片内外的数据搬运开销。40nm CMOS工艺制成的FlexSpIM原型测试结果表明,相比以往固定精度的数字CIM-SNN,比特归一化能效提升2倍,且支持逐位粒度的分辨率重构。在大规模系统中,最高可节省90%能耗,在IBM DVS手势数据集上达到95.8%的先进分类准确率。

原文摘要 · Abstract (English)

Compute-in-memory (CIM) accelerators for spiking neural networks (SNNs) are promising solutions to enable $μ$s-level inference latency and ultra-low energy in edge vision applications. Yet, their current lack of flexibility at both the circuit and system levels prevents their deployment in a wide range of real-life scenarios. In this work, we propose a novel digital CIM macro that supports arbitrary operand resolution and shape, with a unified CIM storage for weights and membrane potentials. These circuit-level techniques enable a hybrid weight- and output-stationary dataflow at the system level to maximize operand reuse, thereby minimizing costly on- and off-chip data movements during the SNN execution. Measurement results of a fabricated FlexSpIM prototype in 40-nm CMOS demonstrate a 2$\times$ increase in bit-normalized energy efficiency compared to prior fixed-precision digital CIM-SNNs, while providing resolution reconfiguration with bitwise granularity. Our approach can save up to 90% energy in large-scale systems, while reaching a state-of-the-art classification accuracy of 95.8% on the IBM DVS gesture dataset.

存内计算脉冲神经网络能效优化边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。