用轻量级芯片压缩脑电数据,提升植入式设备续航与传输效率。
Neural Signal Compression using RAMAN tinyML Accelerator for BCI Applications
- 用卷积自编码器在边缘芯片上实现高达150倍的脑电信号压缩。
- 硬件软件协同优化后参数存储减少32.4%,每通道功耗仅15.1微瓦。
- 适用于高密度脑机接口,适合长期植入设备中的低功耗信号处理。
高质量多通道神经记录对神经科学研究和临床应用至关重要。大规模脑记录常产生海量数据,需无线传输以供后续离线分析与解码,尤其在使用数百至数千电极的高密度皮层记录脑机接口(BCIs)中尤为显著。然而,原始神经数据传输受限于通信带宽,易引发过热问题。为此,我们提出一种基于卷积自编码器(CAEs)的神经信号压缩方案,对局部场电位(LFPs)实现最高150倍的压缩比。将CAE编码器部署于专为边缘计算设计的能效型tinyML加速器RAMAN上。RAMAN通过激活与权重稀疏性,采用零值跳过、门控与权重压缩技术提升能效。同时,采用面向硬件的平衡随机剪枝策略对CAE编码器进行模型剪枝,解决负载不均问题,消除索引开销,使参数存储减少达32.4%。版图后仿真显示,该RAMAN编码器可在TSMC 65nm CMOS工艺下实现,每通道核心面积仅0.0187 mm²。工作频率2 MHz,供电电压1.2 V时,所提DS-CAE1模型每通道功耗为15.1 μW。功能验证方面,已在Efinix Ti60 FPGA上部署,占用37.3k LUTs和8.6k触发器。对两组猴脑神经数据离线重建,信噪失真比(SNDR)分别为22.6 dB和27.4 dB,R²得分分别为0.81和0.94。
原文摘要 · Abstract (English)
High-quality, multi-channel neural recording is indispensable for neuroscience research and clinical applications. Large-scale brain recordings often produce vast amounts of data that must be wirelessly transmitted for subsequent offline analysis and decoding, especially in brain-computer interfaces (BCIs) utilizing high-density intracortical recordings with hundreds or thousands of electrodes. However, transmitting raw neural data presents significant challenges due to limited communication bandwidth and resultant excessive heating. To address this challenge, we propose a neural signal compression scheme utilizing Convolutional Autoencoders (CAEs), which achieves a compression ratio of up to 150 for compressing local field potentials (LFPs). The CAE encoder section is implemented on RAMAN, an energy-efficient tinyML accelerator designed for edge computing. RAMAN leverages sparsity in activation and weights through zero skipping, gating, and weight compression techniques. Additionally, we employ hardware-software co-optimization by pruning the CAE encoder model parameters using a hardware-aware balanced stochastic pruning strategy, resolving workload imbalance issues and eliminating indexing overhead to reduce parameter storage requirements by up to 32.4%. Post layout simulation shows that the RAMAN encoder can be implemented in a TSMC 65-nm CMOS process, occupying a core area of 0.0187 mm2 per channel. Operating at a clock frequency of 2 MHz and a supply voltage of 1.2 V, the estimated power consumption is 15.1 uW per channel for the proposed DS-CAE1 model. For functional validation, the RAMAN encoder was also deployed on an Efinix Ti60 FPGA, utilizing 37.3k LUTs and 8.6k flip-flops. The compressed neural data from RAMAN is reconstructed offline with SNDR of 22.6 dB and 27.4 dB, along with R2 scores of 0.81 and 0.94, respectively, evaluated on two monkey neural recordings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。