无训练深度池化网络在噪声环境下音频事件检测表现优异,适合边缘设备部署。
An Analysis of Untrained Deep Reservoir Networks for Audio Surveillance
- 采用无训练双向池化网络,通过不同深度配置实现音频事件识别。
- 深层模型在低信噪比下更鲁棒,浅层模型效率更高,适合边缘设备。
- 对多种输入特征(如log-Mel、MFCC)均保持稳定性能,适用性广。
本文研究了基于残差计算(RC)范式的无训练递归模型在音频监控中的应用,聚焦于从浅层到深层配置的双向回声状态网络(Bidirectional ESNs),用于紧急声音事件检测。在多类设置下,针对MIVIA Audio Events数据集,评估了不同信噪比(SNR)水平下的模型表现,旨在权衡深度、识别性能与计算效率。对比了全训练的双向长短期记忆网络(BiLSTMs)和卷积-递归神经网络(CRNNs)基线模型。结果表明,深浅级联的池化模型均达到有竞争力的识别率:深层模型在高噪声条件下更鲁棒,浅层模型则具备最优效率,尤其适用于NVIDIA Orin等边缘设备。此外,该方法在不同输入表示(如log-Mel谱图、分辨率可调的MFCC)下仍保持良好鲁棒性。这些发现表明,无训练池化架构是资源受限音频监控场景的有力候选方案。
原文摘要 · Abstract (English)
In this paper, we investigate untrained recurrent models from the Reservoir Computing (RC) paradigm for audio surveillance, focusing on bidirectional Echo State Networks with different depths, from shallow to deep configurations, for emergency sound event detection. We evaluate these models on the MIVIA Audio Events dataset in a multiclass setting across different Signal-to-Noise Ratio (SNR) levels, with the goal of assessing the trade-off between depth, recognition performance, and computational efficiency. We compare the proposed architectures against fully trained recurrent and convolutional-recurrent baselines, namely Bidirectional Long Short-Term Memory networks (BiLSTMs) and Convolutional Recurrent Neural Networks (CRNNs). Results show that deep and shallow reservoir-based models achieve competitive recognition rates, with deeper variants being more robust in highly noisy conditions and shallower ones offering the most favorable efficiency profile, particularly on edge devices such as the NVIDIA Orin. In addition, the proposed approach remains robust across different input representations, including log-Mel spectrograms and MFCCs with varying resolutions. These findings highlight untrained reservoir architectures as a promising solution for resource-constrained audio surveillance scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。