arXiv:2509.13853cs.SDcs.CL2025-09中稿 · ICASSP 2025被引 3

通过特征扰动与噪声监督对比学习,提升异常声音检测准确率

Noise Supervised Contrastive Learning and Feature-Perturbed for Anomalous Sound Detection

  • 引入单阶段噪声监督对比学习,通过嵌入空间特征扰动增强模型区分能力
  • 在DCASE 2020数据集上达95.71% AUC,显著降低同类型设备误报
  • 提出新时频特征TFgram,无需复杂预处理即可捕捉关键异常信息

无监督异常声音检测旨在仅用正常音频训练模型来识别未知异常声音。尽管自监督方法已有进展,但针对不同机器同类样本仍频繁误报的问题尚未解决。本文提出一种新型训练方法——单阶段噪声监督对比学习(OS-SCL),通过在嵌入空间中扰动特征并采用单阶段噪声监督对比学习,有效缓解该问题。在DCASE 2020挑战赛任务2中,仅使用Log-Mel特征即达到94.64% AUC、88.42% pAUC和89.24% mAUC;进一步提出从原始音频提取的时频特征TFgram,最终实现95.71% AUC、90.23% pAUC和91.23% mAUC。源代码已公开于:www.github.com/huangswt/OS-SCL。

原文摘要 · Abstract (English)

Unsupervised anomalous sound detection aims to detect unknown anomalous sounds by training a model using only normal audio data. Despite advancements in self-supervised methods, the issue of frequent false alarms when handling samples of the same type from different machines remains unresolved. This paper introduces a novel training technique called one-stage supervised contrastive learning (OS-SCL), which significantly addresses this problem by perturbing features in the embedding space and employing a one-stage noisy supervised contrastive learning approach. On the DCASE 2020 Challenge Task 2, it achieved 94.64\% AUC, 88.42\% pAUC, and 89.24\% mAUC using only Log-Mel features. Additionally, a time-frequency feature named TFgram is proposed, which is extracted from raw audio. This feature effectively captures critical information for anomalous sound detection, ultimately achieving 95.71\% AUC, 90.23\% pAUC, and 91.23\% mAUC. The source code is available at: \underline{www.github.com/huangswt/OS-SCL}.

异常检测声音分析对比学习特征提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。