用简易可解释的模型实现超低功耗语音关键词识别,功耗仅16.58微瓦。
TsetlinKWS: A 65nm 16.58uW, 0.63mm2 State-Driven Convolutional Tsetlin Machine-Based Accelerator For Keyword Spotting
- 结合频谱特征与卷积,提升卷积型特塞林机在语音任务上的准确率。
- 模型压缩率达9.84倍,推理仅需90.7万次逻辑操作。
- 专为特塞林机设计硬件架构,适合超低功耗边缘设备部署。
特塞林机(TM)因结构简单、可解释性强,被视为神经网络的低功耗替代方案。然而其在语音任务上的表现仍有限。本文提出首个面向12个关键词识别任务的卷积特塞林机(CTM)软硬协同框架TsetlinKWS。首先,引入新型梅尔频率谱系数与谱流(MFSC-SF)特征提取方法,结合谱卷积,使CTM在该任务上首次达到87.35%的竞争力准确率。其次,提出优化分组块压缩行存储(OG-BCSR)算法,实现9.84倍的模型尺寸压缩,显著提升存储效率。最后,设计状态驱动型硬件架构,充分利用数据重用与稀疏性,实现高能效。系统在65 nm工艺下验证,核心面积0.63 mm²,工作电压0.7 V时功耗仅16.58 μW。每次推理仅需907,000次逻辑运算,相比现有最优关键词识别加速器降低10倍,使特塞林机成为超低功耗语音应用的理想候选。
原文摘要 · Abstract (English)
The Tsetlin Machine (TM) has recently attracted attention as a low-power alternative to neural networks due to its simple and interpretable inference mechanisms. However, its performance on speech-related tasks remains limited. This paper proposes TsetlinKWS, the first algorithm-hardware co-design framework for the Convolutional Tsetlin Machine (CTM) on the 12-keyword spotting task. Firstly, we introduce a novel Mel-Frequency Spectral Coefficient and Spectral Flux (MFSC-SF) feature extraction scheme together with spectral convolution, enabling the CTM to reach its first-ever competitive accuracy of 87.35% on the 12-keyword spotting task. Secondly, we develop an Optimized Grouped Block-Compressed Sparse Row (OG-BCSR) algorithm that achieves a remarkable 9.84$\times$ reduction in model size, significantly improving the storage efficiency on CTMs. Finally, we propose a state-driven architecture tailored for the CTM, which simultaneously exploits data reuse and sparsity to achieve high energy efficiency. The full system is evaluated in 65 nm process technology, consuming 16.58 $μ$W at 0.7 V with a compact 0.63 mm$^2$ core area. TsetlinKWS requires only 907k logic operations per inference, representing a 10$\times$ reduction compared to the state-of-the-art KWS accelerators, positioning the CTM as a highly-efficient candidate for ultra-low-power speech applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。