arXiv:2409.03610eess.AS2024-09中稿 · ICASSP 2024被引 15

通过频时双通道学习,提升机器异常声音检测精度

A Dual-Path Framework with Frequency-and-Time Excited Network for Anomalous Sound Detection

  • 设计频时激励网络,同时捕捉声谱的频率与时间特征
  • 在DCASE 2023数据集上达到当前最优性能
  • 适合关注工业声音异常检测的研究者

与人类语音不同,同类型机器生成的声音通常具有稳定的频率特性与明显的时序周期性。然而,将这一双重属性用于异常检测仍研究较少。本文提出一种自动化双路径框架,针对多种机器类型学习显著的频率与时间模式。一条路径采用新型频时激励网络(FTE-Net),从声谱的频率与时间轴中提取关键特征,包含频时分块编码器(FTC-Encoder)与激励网络;另一条路径使用1D卷积网络处理语句级频谱。在DCASE 2023任务2数据集上的实验表明,所提方法达到当前最优性能。此外,通过可视化激励网络的中间特征图,验证了方法的有效性。

原文摘要 · Abstract (English)

In contrast to human speech, machine-generated sounds of the same type often exhibit consistent frequency characteristics and discernible temporal periodicity. However, leveraging these dual attributes in anomaly detection remains relatively under-explored. In this paper, we propose an automated dual-path framework that learns prominent frequency and temporal patterns for diverse machine types. One pathway uses a novel Frequency-and-Time Excited Network (FTE-Net) to learn the salient features across frequency and time axes of the spectrogram. It incorporates a Frequency-and-Time Chunkwise Encoder (FTC-Encoder) and an excitation network. The other pathway uses a 1D convolutional network for utterance-level spectrum. Experimental results on the DCASE 2023 task 2 dataset show the state-of-the-art performance of our proposed method. Moreover, visualizations of the intermediate feature maps in the excitation network are provided to illustrate the effectiveness of our method.

异常检测声音分析深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。