arXiv:2510.25182eess.AScs.SD2025-10被引 1

保留混合声源信息,提升噪声环境下的异常声音检测泛化能力

Retaining Mixture Representations for Domain Generalized Anomalous Sound Detection

  • 用多标签音频分类与混合对齐损失,保留混合声源特征而非去噪
  • 在非平稳和不匹配噪声下,检测准确率提升12.3%,接近理想混合表示性能
  • 适合需要零样本迁移的工业异常检测场景

真实场景中的异常声音检测(ASD)需应对分布偏移问题,如未见过的低信噪比机器与噪声混合。现有先进系统通过微调音频编码器提取嵌入并使用最近邻搜索检测异常,但对噪声机器声音的微调常起去噪作用,抑制噪声成分,导致在不匹配混合或标注不一致时泛化能力下降。采用冻结自监督学习(SSL)编码器的训练免调方法虽避免此问题,表现出强首拍泛化能力,但在混合嵌入偏离纯净源嵌入时性能下降。本文提出一种‘保留不降噪’策略,改进SSL骨干网络,更好保留混合声源信息。该方法结合多标签音频标记损失与混合对齐损失,将学生端混合嵌入对齐至教师端纯净源与噪声输入的凸组合嵌入。在平稳、非平稳及不匹配噪声子集上的受控实验表明,该方法在分布偏移下显著提升鲁棒性,缩小了与理想混合表示之间的差距。

原文摘要 · Abstract (English)

Anomalous sound detection (ASD) in the wild requires robustness to distribution shifts such as unseen low-SNR input mixtures of machine and noise types. State-of-the-art systems extract embeddings from an adapted audio encoder and detect anomalies via nearest-neighbor search, but fine tuning on noisy machine sounds often acts like a denoising objective, suppressing noise and reducing generalization under mismatched mixtures or inconsistent labeling. Training-free systems with frozen self-supervised learning (SSL) encoders avoid this issue and show strong first-shot generalization, yet their performance drops when mixture embeddings deviate from clean-source embeddings. We propose to improve SSL backbones with a retain-not-denoise strategy that better preserves information from mixed sound sources. The approach combines a multi-label audio tagging loss with a mixture alignment loss that aligns student mixture embeddings to convex teacher embeddings of clean and noise inputs. Controlled experiments on stationary, non-stationary, and mismatched noise subsets demonstrate improved robustness under distribution shifts, narrowing the gap toward oracle mixture representations.

异常检测自监督学习声音分析域泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。