用低秩微调提升预训练音频模型的异常声音检测能力
Improving Anomalous Sound Detection via Low-Rank Adaptation Fine-Tuning of Pre-Trained Audio Models
- 采用低秩适配(LoRA)微调预训练音频模型,降低数据依赖
- 在DCASE2023数据集上达到77.75%准确率,比SOTA提升6.48%
- 适合工业场景中数据稀缺下的异常声音检测任务
异常声音检测(ASD)在工业环境中因人工智能技术的应用而受到广泛关注。然而,由于数据采集困难和环境因素复杂,现有ASD系统难以直接部署。本文提出一种基于预训练音频模型的鲁棒ASD方法,利用设备运行数据进行微调,并采用SpecAug作为数据增强策略。同时,研究了使用低秩适配(LoRA)微调替代全量微调的效果,以应对微调数据有限的问题。在DCASE2023 Task 2数据集上的实验表明,该方法在评估集上达到77.75%的准确率,较此前最优模型(包括顶级卷积网络与语音预训练模型)提升6.48%,验证了预训练音频模型结合LoRA微调的有效性。消融实验进一步证明了所提方案的优越性。
原文摘要 · Abstract (English)
Anomalous Sound Detection (ASD) has gained significant interest through the application of various Artificial Intelligence (AI) technologies in industrial settings. Though possessing great potential, ASD systems can hardly be readily deployed in real production sites due to the generalization problem, which is primarily caused by the difficulty of data collection and the complexity of environmental factors. This paper introduces a robust ASD model that leverages audio pre-trained models. Specifically, we fine-tune these models using machine operation data, employing SpecAug as a data augmentation strategy. Additionally, we investigate the impact of utilizing Low-Rank Adaptation (LoRA) tuning instead of full fine-tuning to address the problem of limited data for fine-tuning. Our experiments on the DCASE2023 Task 2 dataset establish a new benchmark of 77.75% on the evaluation set, with a significant improvement of 6.48% compared with previous state-of-the-art (SOTA) models, including top-tier traditional convolutional networks and speech pre-trained models, which demonstrates the effectiveness of audio pre-trained models with LoRA tuning. Ablation studies are also conducted to showcase the efficacy of the proposed scheme.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。