arXiv:2509.08717cs.SDcs.AI2025-09中稿 · IEEE ICTAI 2025被引 7

用多种可解释性方法分析鸟类叫声分类模型,提升AI决策可信度。

Explainability of CNN Based Classification Models for Acoustic Signal

  • 结合多种XAI技术解析CNN对鸟鸣声谱图的分类依据。
  • 模型分类准确率达94.8%,多方法解释提升决策透明度。
  • 适合关注生物声学、AI可解释性的研究者参考。

可解释人工智能(XAI)已成为解读复杂深度学习模型预测结果的关键工具。尽管XAI已在多个声学领域得到应用,但在生物声学——即分析生物体发出的音频信号——中的使用仍相对不足。本文研究了北美洲范围内具有显著地理变异的鸟类物种的鸣叫声。将音频记录转化为声谱图图像,并用于训练深度卷积神经网络(CNN)进行分类,取得了94.8%的准确率。为解释模型预测,我们应用了两类XAI技术:模型无关的(LIME、SHAP)和模型特定的(DeepLIFT、Grad-CAM)。这些方法生成了不同但互补的解释,综合使用时能提供更完整、更可理解的模型决策洞察。本研究强调,结合多种XAI技术有助于提升信任度与跨领域适用性,不仅适用于声学信号分析,也具广泛应用于特定领域任务的潜力。

原文摘要 · Abstract (English)

Explainable Artificial Intelligence (XAI) has emerged as a critical tool for interpreting the predictions of complex deep learning models. While XAI has been increasingly applied in various domains within acoustics, its use in bioacoustics, which involves analyzing audio signals from living organisms, remains relatively underexplored. In this paper, we investigate the vocalizations of a bird species with strong geographic variation throughout its range in North America. Audio recordings were converted into spectrogram images and used to train a deep Convolutional Neural Network (CNN) for classification, achieving an accuracy of 94.8\%. To interpret the model's predictions, we applied both model-agnostic (LIME, SHAP) and model-specific (DeepLIFT, Grad-CAM) XAI techniques. These techniques produced different but complementary explanations, and when their explanations were considered together, they provided more complete and interpretable insights into the model's decision-making. This work highlights the importance of using a combination of XAI techniques to improve trust and interoperability, not only in broader acoustics signal analysis but also argues for broader applicability in different domain specific tasks.

可解释AI生物声学深度学习分类模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。