arXiv:2601.12660cs.SDcs.LG2026-01中稿 · the IEEE Internati…

对比自编码器与掩码自编码器,提升音频异常检测的解释可信度。

Toward Faithful Explanations in Acoustic Anomaly Detection

  • 用掩码自编码器增强模型对异常区域的定位精度。
  • 在真实工业场景中,掩码模型解释更贴近真实异常时间点。
  • 提出扰动评估法验证解释区域的有效性,适合工业级可解释系统设计。

可解释性对用户信任真实世界异常检测至关重要。尽管深度学习模型表现优异,但缺乏透明性。本文研究基于自编码器的音频异常检测模型的可解释性,比较标准自编码器(AE)与掩码自编码器(MAE)在检测性能与可解释性上的差异。采用误差图、显著性图、SmoothGrad、Integrated Gradients、GradSHAP 和 Grad-CAM 等多种归因方法。结果显示,虽然 MAE 检测性能略低,但其解释始终更忠实且时间精度更高,更贴近真实异常。为评估解释区域的相关性,提出一种基于扰动的忠实度度量:将解释区域替换为重建结果以模拟正常输入。实验基于真实工业场景,表明在异常检测流程中融入可解释性至关重要,且掩码训练可在不牺牲性能前提下显著提升解释质量。

原文摘要 · Abstract (English)

Interpretability is essential for user trust in real-world anomaly detection applications. However, deep learning models, despite their strong performance, often lack transparency. In this work, we study the interpretability of autoencoder-based models for audio anomaly detection, by comparing a standard autoencoder (AE) with a mask autoencoder (MAE) in terms of detection performance and interpretability. We applied several attribution methods, including error maps, saliency maps, SmoothGrad, Integrated Gradients, GradSHAP, and Grad-CAM. Although MAE shows a slightly lower detection, it consistently provides more faithful and temporally precise explanations, suggesting a better alignment with true anomalies. To assess the relevance of the regions highlighted by the explanation method, we propose a perturbation-based faithfulness metric that replaces them with their reconstructions to simulate normal input. Our findings, based on experiments in a real industrial scenario, highlight the importance of incorporating interpretability into anomaly detection pipelines and show that masked training improves explanation quality without compromising performance.

异常检测可解释性音频分析自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。