自适应语音转脉冲编码让神经形态语音识别更高效准确
Adaptive Speech-to-Spike Encoding for Spiking Neural Networks

- 设计可学习的残差编码器,与脉冲网络端到端联合训练
- 在GSC-v2上达94.97%准确率,35k参数版本仍超89.8%
- 编码器专注任务特征而非信号还原,适合硬件部署
连续语音信号与离散事件驱动处理之间的不匹配仍是神经形态语音处理的核心瓶颈。现有系统多依赖固定脉冲编码器,迫使下游脉冲神经网络(SNN)自行补偿非自适应输入表示。为此,我们提出一种可学习的残差语音-脉冲编码器,与递归漏电积分-放电(R-LIF)主干网络联合端到端训练。在Google语音命令v2(GSC-v2)基准上,该方法最高实现94.97%准确率。值得注意的是,紧凑的35,000参数变体达到89.8%准确率,与此前需多一个数量级参数的基线相当或更优。通过编码器聚焦分析(包括线性探测与梯度残差检查),发现编码器不追求信号忠实重建,而是学习任务对齐的脉冲表示以增强类别可分性。最后,在相同架构与训练条件下,对比生物启发、硬件友好的信用分配机制——直接反馈对齐(DFA)与代理梯度反向传播(BPTT)。结果表明,DFA达到91.5%准确率,量化了生物启发学习规则在现代神经形态音频任务中的性能权衡。
原文摘要 · Abstract (English)
The mismatch between continuous acoustic signals and discrete event-driven processing remains a fundamental bottleneck for neuromorphic speech processing. Current systems typically rely on fixed spike encoders, forcing downstream Spiking Neural Networks (SNNs) to compensate for non-adaptive input representations. To address this, we present a learnable residual speech-to-spike encoder jointly trained end-to-end with a Recurrent Leaky Integrate-and-Fire (R-LIF) backbone. We validate this approach on the Google Speech Commands v2 (GSC-v2) benchmark, achieving up to 94.97% accuracy. Notably, the learned encoder remains highly parameter-efficient with a compact 35k-parameter variant that reaches 89.8%, matching or exceeding prior baselines that require an order of magnitude more parameters. Our encoder-focused analysis, including linear probing and gradient-residual inspection, indicates that the encoder does not target faithful signal reconstruction but instead learns task-aligned spike representations that enhance class separability. Finally, we benchmark bio-inspired, hardware-friendly credit assignment by comparing Direct Feedback Alignment (DFA) with surrogate-gradient BPTT under identical architectures and training conditions. We find that DFA reaches 91.5% accuracy, quantifying the performance trade-off of bio-inspired learning rules for modern neuromorphic audio.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。