arXiv:2606.22901eess.AScs.AI2026-06被引 1

提出新方法评估语音识别模型的注意力机制,让AI决策更透明。

Explainable AI in Speaker Recognition -- Attention Map Visualisation and Evaluation

论文配图:Explainable AI in Speaker Recognition -- Attention Map Visualisation and Evaluation
图 1 · 摘自论文原文
  • 改进已有评估方法,提出修正版随机输入采样评估算法
  • 实测发现两种注意力可视化方法在不同场景各有优势
  • 适合关注AI可解释性与语音识别模型可信度的研究者

解释人工智能系统(尤其是神经网络)的决策过程是可解释人工智能(XAI)的核心。类比人类注意力机制,神经网络被认为具备选择性处理信息的能力。本文聚焦于分析与可视化神经网络的注意力机制,实验基于训练用于从语音片段中识别说话人身份的神经网络。以往研究多采用基于类激活图(CAM)的方法生成注意力图,每个输入对应一张图,指示模型决策时关注的区域。然而,这些注意力图的评估仍缺乏系统研究。本文系统回顾现有注意力图评估算法,厘清关键概念并指出其不足,进而提出改进版本——修正版随机输入采样解释评估算法(Modified RISE-eval)。利用该算法,对两种代表性CAM方法(GradCAM与LayerCAM)在特定说话人识别网络上的注意力图进行评估。结果表明,两种方法在不同实验条件下各具优势。

原文摘要 · Abstract (English)

Explaining and understanding the decision-making process of artificial intelligence (AI) systems, particularly those implemented by neural networks, falls within the field of explainable AI (XAI). Analogous to the human attention mechanism, neural networks are assumed to possess their own attention mechanisms that selectively process information during decision-making. This work proposes to study one XAI topic: analysing and visualising the attention mechanisms of neural networks. Our experiments are performed on speaker recognition neural networks that are trained to identify speaker identity from a given utterance. Previous studies have widely used class activation map (CAM)-based methods to analyse and visualise the attention mechanisms of neural networks. Each of these methods produces an attention map for each network input, highlighting which input regions are selectively processed when the speaker recognition network makes decisions. However, the evaluation of attention maps produced by these methods remains largely underexplored. This work systematically reviews an existing attention map evaluation algorithm, establishing key concepts and identifying its shortcomings. On the basis of this existing evaluation algorithm, a new version is then proposed to address the identified shortcomings, called the Modified Randomised Input Sampling for Explanation - Evaluation algorithm (Modified RISE-eval). Using Modified RISE-eval, we evaluate the attention maps produced by two representative CAM-based methods, GradCAM and LayerCAM, applied to a certain speaker recognition network. The evaluation results demonstrate that GradCAM and LayerCAM each exhibit distinct advantages when applied under different experimental conditions in the speaker recognition task.

可解释AI注意力机制语音识别评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。