用脉冲神经网络设计低功耗音频事件触发器,显著降低计算开销。
A Neuromorphic Trigger for Efficient Audio Event Detection

- 用脉冲神经网络构建轻量级触发器,仅处理关键音频片段。
- 在ASD任务中达到0.97的F1分数,在SED任务中减少42.6倍运算量。
- 适合实时、低功耗设备上的音频事件检测应用。
连续音频流的高效处理仍是实时和资源受限系统的关键挑战。本文提出一种基于脉冲神经网络(SNN)的类脑触发器,用于音频事件检测,可选择性地将输入传递给下游模型。该触发器作为灵活且低成本的前端,识别重要音频段,使更复杂的模型仅处理这些片段以完成分类等任务。触发器采用轻量级全连接SNN结构,并结合闭-开滤波器进行后处理。在两个代表性任务上进行评估:异常声音检测(ASD)和声音事件检测(SED)。在类无关的URBAN-SED数据集上,一秒钟片段级别的F1得分为0.97,表明其在识别相关音频区域方面具有高可靠性。在DCASE 2017 Challenge Task 2数据集上,与Dang分类器结合使用时,计算量降低42.6倍,事件误差率下限从0.41降至0.25。结果表明,类脑触发器可作为实时、节能的前端过滤器,大幅降低计算成本。
原文摘要 · Abstract (English)
Efficient processing of continuous audio streams remains a key challenge for real-time and resource-constrained systems. This paper introduces a neuromorphic trigger for audio event detection, based on a spiking neural network (SNN) that selectively gates input to downstream models. The proposed neuromorphic trigger acts as a flexible low-cost front-end, identifying salient audio segments and enabling these to be processed by a more computationally intensive model for tasks such as classification. The trigger is implemented as a lightweight fully connected SNN using a close-open filter for postprocessing, and is evaluated on two representative tasks: Anomalous Sound Detection (ASD) and Sound Event Detection (SED). For ASD, the trigger achieves a one-second segment-based F1 score of 0.97 on a class-agnostic form of the URBAN-SED dataset, demonstrating high reliability in identifying relevant audio regions. For SED, the trigger is combined with the Dang classifier on the DCASE 2017 Challenge Task 2 dataset, showing a potential $42.6\times$ reduction in FLOPs while reducing the lower bound of the event-based error rate from 0.41 to 0.25. These results highlight the potential of neuromorphic triggers as real-time, energy-efficient front-end filters, enabling substantial reductions in computational cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。