arXiv:2508.12230cs.SDeess.AS2025-08中稿 · TASLP被引 14

用自监督音频模型提升设备异常声音检测的泛化能力

Exploring Self-Supervised Audio Models for Generalized Anomalous Sound Detection

  • 基于大规模音频预训练模型,结合低秩适配与分组适配模块
  • 在多个基准数据集上显著提升异常检测准确率,尤其在小样本场景
  • 适合工业质检、设备运维等需跨设备泛化的场景

机器异常声音检测(ASD)在多个领域具有重要应用价值,但其泛化性能常受限于数据采集困难和声学环境复杂。受大模型在多领域成功启发,本文提出一种鲁棒的ASD模型,利用在大规模语音与音频数据上预训练的自监督模型。尽管预训练数据与ASD任务存在差异,实验表明预训练仍带来显著收益。为缓解小样本微调中的过拟合问题,探索使用全连接低秩适配(LoRA)替代全量微调。同时提出机器感知分组适配模块,统一框架内捕捉不同机器间的差异,增强泛化性。针对属性标签缺失问题,设计新型目标函数,通过向量量化动态聚类无标签数据,并以双层对比学习优化。方法在包括DCASE 2020-2024五项挑战赛在内的所有基准数据集上评估,结果表明新方法显著优于基线,验证了所提策略的有效性。

原文摘要 · Abstract (English)

Machine anomalous sound detection (ASD) is a valuable technique across various applications. However, its generalization performance is often limited due to challenges in data collection and the complexity of acoustic environments. Inspired by the success of large pre-trained models in numerous fields, this paper introduces a robust ASD model that leverages self-supervised pre-trained models trained on large-scale speech and audio datasets. Although there are inconsistencies between the pre-training datasets and the ASD task, our findings indicate that pre-training still provides substantial benefits for ASD. To mitigate overfitting and retain learned knowledge when fine-tuning with limited data, we explore Fully-Connected Low-Rank Adaptation (LoRA) as an alternative to full fine-tuning. Additionally, we propose a Machine-aware Group Adapter module, which enables the model to capture differences between various machines within a unified framework, thereby enhancing the generalization performance of ASD systems. To address the challenge of missing attribute labels, we design a novel objective function that dynamically clusters unattributed data using vector quantization and optimizes through a dual-level contrastive learning loss. The proposed methods are evaluated on all benchmark datasets, including the DCASE 2020-2024 five ASD challenges, and the experimental results show significant improvements of our new approach and demonstrate the effectiveness of our proposed strategies.

异常检测自监督学习音频分析工业智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。