arXiv:2503.21337cs.ARcs.AI2025-03被引 4

超低功耗语音识别芯片,仅71.2μW实现实时推理。

A 71.2-$μ$W Speech Recognition Accelerator with Recurrent Spiking Neural Network

  • 用稀疏脉冲神经网络+硬件优化,降低计算复杂度
  • 模型压缩96.42%至0.1MB,功耗仅71.2μW
  • 适合可穿戴设备等超低功耗边缘场景

本文提出一种面向边缘设备实时应用的超低功耗语音识别加速器,功耗仅为71.2μW。通过算法与硬件协同优化,设计了一种包含两层循环结构、一层全连接层、且时间步长极低(1或2)的紧凑型脉冲神经网络。2.79-MB模型经剪枝与4位定点量化后,体积缩小96.42%,降至0.1MB。硬件层面采用混合级剪枝、零跳过和合并脉冲技术,将计算复杂度降低90.49%,达13.86 MMAC/S。并行时间步执行缓解了时间步间数据依赖,通过权重共享节省缓冲区功耗。利用脉冲活动稀疏性,输入广播机制消除零计算,进一步节能。该设计基于台积电28nm工艺实现,在100 kHz下实时运行,500 MHz时能效达28.41 TOPS/W,面积效率为1903.11 GOPS/mm²,优于现有先进方案。

原文摘要 · Abstract (English)

This paper introduces a 71.2-$μ$W speech recognition accelerator designed for edge devices' real-time applications, emphasizing an ultra low power design. Achieved through algorithm and hardware co-optimizations, we propose a compact recurrent spiking neural network with two recurrent layers, one fully connected layer, and a low time step (1 or 2). The 2.79-MB model undergoes pruning and 4-bit fixed-point quantization, shrinking it by 96.42\% to 0.1 MB. On the hardware front, we take advantage of \textit{mixed-level pruning}, \textit{zero-skipping} and \textit{merged spike} techniques, reducing complexity by 90.49\% to 13.86 MMAC/S. The \textit{parallel time-step execution} addresses inter-time-step data dependencies and enables weight buffer power savings through weight sharing. Capitalizing on the sparse spike activity, an input broadcasting scheme eliminates zero computations, further saving power. Implemented on the TSMC 28-nm process, the design operates in real time at 100 kHz, consuming 71.2 $μ$W, surpassing state-of-the-art designs. At 500 MHz, it has 28.41 TOPS/W and 1903.11 GOPS/mm$^2$ in energy and area efficiency, respectively.

语音识别脉冲神经网络低功耗边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。