对比三种脉冲编码方法,发现TAE在环境音识别中效果最佳且更省电。
Spike Encoding for Environmental Sound: A Comparative Benchmark
- 采用多频带分析比较TAE、SF、MW三种脉冲编码方法。
- TAE在所有数据集上重建质量最优,脉冲发放率最低。
- 适合需要低功耗处理环境声音的神经形态系统应用。
脉冲神经网络(SNN)具备能效优势,适用于边缘计算场景,但常规传感器数据需先转换为脉冲序列才能进行神经形态处理。环境声音(如城市声景)因频率变化大、背景噪声强及声事件重叠而具有挑战性,而现有基于脉冲的音频编码研究多集中于语音。本文在ESC10、UrbanSound8K和TAU Urban Acoustic Scenes三个数据集上,对比了阈值自适应编码(TAE)、逐步前移(SF)和滑动窗口(MW)三种编码方法。多频带分析表明,TAE在各频段及各类别上的重建质量均优于SF与MW;同时,其脉冲发放率最低,体现更高能效。在标准SNN下游环境声音分类任务中,TAE亦表现最佳。本工作为神经形态环境声音处理中的脉冲编码选择提供了基础性洞察与基准参考。
原文摘要 · Abstract (English)
Spiking Neural Networks (SNNs) offer energy efficient processing suitable for edge applications, but conventional sensor data must first be converted into spike trains for neuromorphic processing. Environmental sound, including urban soundscapes, poses challenges due to variable frequencies, background noise, and overlapping acoustic events, while most spike based audio encoding research has focused on speech. This paper analyzes three spike encoding methods, Threshold Adaptive Encoding (TAE), Step Forward (SF), and Moving Window (MW) across three datasets: ESC10, UrbanSound8K, and TAU Urban Acoustic Scenes. Our multiband analysis shows that TAE consistently outperforms SF and MW in reconstruction quality, both per frequency band and per class across datasets. Moreover, TAE yields the lowest spike firing rates, indicating superior energy efficiency. For downstream environmental sound classification with a standard SNN, TAE also achieves the best performance among the compared encoders. Overall, this work provides foundational insights and a comparative benchmark to guide the selection of spike encoders for neuromorphic environmental sound processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。