用声音分离模型提升异常声音检测的表示学习能力
Representational learning for an anomalous sound detection system with source separation model
- 通过源分离模型分离目标与非目标机器声音,增强表征学习
- 在少量目标数据下仍优于传统自编码器和信号分离方法
- 非目标数据越多,表征能力越强,适合工业异常检测场景
机械设备运行中的异常声音检测因难以泛化异常声学模式而面临挑战。该任务通常被视为无监督学习或新奇检测问题,因全面获取异常声学数据困难。传统方法多采用自编码器架构或带辅助任务的表示学习,但各有局限:自编码器仅使用目标机器运行声音,而辅助任务虽可引入多样化声学输入,其表征可能与异常特征无关。本文提出基于源分离模型(CMGAN)的训练方法,旨在从目标与非目标声音混合信号中分离出非目标声音。该方法可有效利用多样化的机器声音,支持复杂神经网络在小样本下的训练。实验表明,所提方法在性能上优于传统自编码器及聚焦目标信号分离的方法。此外,当非目标数据量增加时,表征学习能力持续提升,即使目标数据量保持不变。
原文摘要 · Abstract (English)
The detection of anomalous sounds in machinery operation presents a significant challenge due to the difficulty in generalizing anomalous acoustic patterns. This task is typically approached as an unsupervised learning or novelty detection problem, given the complexities associated with the acquisition of comprehensive anomalous acoustic data. Conventional methodologies for training anomalous sound detection systems primarily employ auto-encoder architectures or representational learning with auxiliary tasks. However, both approaches have inherent limitations. Auto-encoder structures are constrained to utilizing only the target machine's operational sounds, while training with auxiliary tasks, although capable of incorporating diverse acoustic inputs, may yield representations that lack correlation with the characteristic acoustic signatures of anomalous conditions. We propose a training method based on the source separation model (CMGAN) that aims to isolate non-target machine sounds from a mixture of target and non-target class acoustic signals. This approach enables the effective utilization of diverse machine sounds and facilitates the training of complex neural network architectures with limited sample sizes. Our experimental results demonstrate that the proposed method yields better performance compared to both conventional auto-encoder training approaches and source separation techniques that focus on isolating target machine signals. Moreover, our experimental results demonstrate that the proposed method exhibits the potential for enhanced representation learning as the quantity of non-target data increases, even while maintaining a constant volume of target class data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。