用噪声感知自监督学习提升机器异常声音检测精度
Anomalous Sound Detection Meets Noise-Aware Self-Supervised Learning
- 利用远距离麦克风噪声信号辅助提取近距信号的纯净表征
- 在DCASE 2026挑战赛中,NA-BEATs系统以70.24%得分夺冠
- 适用于工业设备巡检等存在复杂噪声环境的异常检测场景
本文提出噪声感知自监督学习(NA-SSL)模型用于噪声感知异常声音检测(NA-ASD)。该任务采用双麦克风录音:一个靠近目标机器,另一个远离以捕捉背景噪声。我们通过多样化音频数据集模拟双通道录音,训练NA-SSL模型,利用远距麦克风主导的噪声信号作为辅助信息,提取近距麦克风信号的干净自监督表征。这些模型作为前端嵌入标准异常声音检测框架。在DCASE 2026挑战赛任务2开发数据集上的实验表明,该框架在三种基础自监督模型(BEATs、EAT和Dasheng)上均有效,无论是否进行判别式微调。挑战赛结果进一步验证了方法的有效性:所提NA-BEATs系统以70.24%的官方得分大幅领先第二名的65.46%。
原文摘要 · Abstract (English)
In this paper, we introduce noise-aware self-supervised learning (NA-SSL) models for noise-aware anomalous sound detection (NA-ASD). NA-ASD is an ASD task with two-channel audio recordings, where one microphone is located close to the target machine and the other is located farther away to capture noise. For this task, we simulate two-channel recordings using diverse audio datasets and train NA-SSL models to extract clean SSL representations of the close-microphone signal by using the far-microphone recording dominated by background noise as auxiliary information. The NA-SSL models are then used as frontends in the standard ASD framework. Our experimental evaluation on the DCASE 2026 Challenge Task 2 development dataset demonstrates the effectiveness of the NA-SSL framework across three base SSL models (BEATs, EAT, and Dasheng), both with and without discriminative fine-tuning. Furthermore, the challenge results proved the effectiveness of the proposed approach, where the NA-BEATs system won the challenge by a large margin, achieving an official score of 70.24%, while the second-place system achieved 65.46%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。