用视觉算法检测音频异常,定位更准且可解释。
From Vision to Sound: Advancing Audio Anomaly Detection with Vision-Based Algorithms
- 借鉴视觉异常检测思路,用预训练模型提取音频特征
- 实现频谱图中异常的精细时空定位,提升可解释性
- 适合工业与环境场景,对异常位置和时间精准识别
视觉异常检测(VAD)近年发展出利用预训练特征提取器生成嵌入的先进算法。受此启发,我们探索将此类方法迁移至音频领域,以解决音频异常检测(AAD)问题。与多数现有方法仅分类异常不同,本方法可在频谱图中实现异常的细粒度时频定位,显著提升结果的可解释性,使用户更清晰了解异常发生的时间与位置,增强实用性。我们在工业与环境基准数据集上验证了该方法的有效性,证明了VAD技术在音频信号异常检测中的适用性,并通过局部化识别提升了系统的可解释性与实际应用价值。
原文摘要 · Abstract (English)
Recent advances in Visual Anomaly Detection (VAD) have introduced sophisticated algorithms leveraging embeddings generated by pre-trained feature extractors. Inspired by these developments, we investigate the adaptation of such algorithms to the audio domain to address the problem of Audio Anomaly Detection (AAD). Unlike most existing AAD methods, which primarily classify anomalous samples, our approach introduces fine-grained temporal-frequency localization of anomalies within the spectrogram, significantly improving explainability. This capability enables a more precise understanding of where and when anomalies occur, making the results more actionable for end users. We evaluate our approach on industrial and environmental benchmarks, demonstrating the effectiveness of VAD techniques in detecting anomalies in audio signals. Moreover, they improve explainability by enabling localized anomaly identification, making audio anomaly detection systems more interpretable and practical.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。