arXiv:2409.08587eess.AScs.SD2024-09中稿 · paper: Workshop on…被引 5

用频率追踪特征提升小数据下警报声识别效果

Frequency Tracking Features for Data-Efficient Deep Siren Identification

  • 基于单参数自适应陷波器追踪警报声频率变化
  • 小样本训练下性能优于传统频谱模型,准确率更高
  • 模型轻量、泛化强,适合资源受限场景

城市声景中识别警报声对智能车辆安全至关重要。现有神经网络方法在充足数据下表现优异,但实际中数据有限,且需应对未见声学环境。本文针对警报声具有周期性基频变化的谐波特性,提出一种基于单参数自适应陷波器的低复杂度频率追踪特征提取方法。该特征用于设计小型卷积网络,可在有限数据下高效训练。实验表明,该模型在小样本条件下持续优于传统频谱基模型,具备更强跨域泛化能力,且模型尺寸更小。

原文摘要 · Abstract (English)

The identification of siren sounds in urban soundscapes is a crucial safety aspect for smart vehicles and has been widely addressed by means of neural networks that ensure robustness to both the diversity of siren signals and the strong and unstructured background noise characterizing traffic. Convolutional neural networks analyzing spectrogram features of incoming signals achieve state-of-the-art performance when enough training data capturing the diversity of the target acoustic scenes is available. In practice, data is usually limited and algorithms should be robust to adapt to unseen acoustic conditions without requiring extensive datasets for re-training. In this work, given the harmonic nature of siren signals, characterized by a periodically evolving fundamental frequency, we propose a low-complexity feature extraction method based on frequency tracking using a single-parameter adaptive notch filter. The features are then used to design a small-scale convolutional network suitable for training with limited data. The evaluation results indicate that the proposed model consistently outperforms the traditional spectrogram-based model when limited training data is available, achieves better cross-domain generalization and has a smaller size.

音频识别小样本学习特征提取警报检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。