arXiv:2410.22056eess.AScs.SD2024-10被引 1

无需训练即可检测异常声音并生成相关原因描述

Retrieval-Augmented Approach for Unsupervised Anomalous Sound Detection and Captioning without Model Training

  • 利用预训练的CLAP模型嵌入空间进行检索式原因生成
  • 检测结果与生成描述高度一致,无需额外标注数据
  • 适合需要可解释性的工业故障检测场景

本文提出一种无监督异常声音检测(UASD)及原因描述生成方法。现有方法虽可对正常与异常声音对生成差异描述,但需独立训练且依赖大量差异标签,导致描述与检测结果不一致。所提方法基于预训练的CLAP模型,在其嵌入空间中采用检索增强策略生成异常声音的原因描述,实现检测与描述的一致性,无需模型训练。主观评价与逐样本分析实验验证了该方法的有效性。

原文摘要 · Abstract (English)

This paper proposes a method for unsupervised anomalous sound detection (UASD) and captioning the reason for detection. While there is a method that captions the difference between given normal and anomalous sound pairs, it is assumed to be trained and used separately from the UASD model. Therefore, the obtained caption can be irrelevant to the differences that the UASD model captured. In addition, it requires many caption labels representing differences between anomalous and normal sounds for model training. The proposed method employs a retrieval-augmented approach for captioning of anomalous sounds. Difference captioning in the embedding space output by the pre-trained CLAP (contrastive language-audio pre-training) model makes the anomalous sound detection results consistent with the captions and does not require training. Experiments based on subjective evaluation and a sample-wise analysis of the output captions demonstrate the effectiveness of the proposed method.

异常检测音频理解零样本可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。