无需异常音样本训练,通过音色特征差异解释机器异常声音
Timbre Difference Capturing in Anomalous Sound Detection
- 用预定义音色属性替代自由文本描述,避免依赖异常音频训练
- 结合k近邻算法在音频嵌入空间中同时实现检测与音色差异估计
- 基于心理声学模型量化音色变化,适合工业设备状态监测场景
本文提出一种用于解释异常声音差异的框架,旨在解决异常声音检测(ASD)中难以获取异常音频样本的问题。现有方法需依赖异常声音进行训练,但在实际设备监测中此类数据往往不可得。为此,本文提出不依赖异常样本的新型解释策略:聚焦于预定义音色属性的差异,利用心理声学研究构建的音色模型计算客观指标,从而无须训练机器学习模型即可估计正常声音中音色属性的变化。为应对正常数据多样性带来的干扰,进一步设计了一种联合方法,在音频嵌入空间中基于k近邻实现异常检测与音色差异估计。在MIMII DG数据集上的实验验证了该方法的有效性。
原文摘要 · Abstract (English)
This paper proposes a framework of explaining anomalous machine sounds in the context of anomalous sound detection~(ASD). While ASD has been extensively explored, identifying how anomalous sounds differ from normal sounds is also beneficial for machine condition monitoring. However, existing sound difference captioning methods require anomalous sounds for training, which is impractical in typical machine condition monitoring settings where such sounds are unavailable. To solve this issue, we propose a new strategy for explaining anomalous differences that does not require anomalous sounds for training. Specifically, we introduce a framework that explains differences in predefined timbre attributes instead of using free-form text captions. Objective metrics of timbre attributes can be computed using timbral models developed through psycho-acoustical research, enabling the estimation of how and what timbre attributes have changed from normal sounds without training machine learning models. Additionally, to accurately determine timbre differences regardless of variations in normal training data, we developed a method that jointly conducts anomalous sound detection and timbre difference estimation based on a k-nearest neighbors method in an audio embedding space. Evaluation using the MIMII DG dataset demonstrated the effectiveness of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。