用熵引导注意力提升语音模型解释性,让预测更透明可信。
Listening with Attention: Entropy-Guided Explainability for Transformer-Based Audio Models

- 基于熵引导的注意力加权,定位关键模型头和层
- 忠实度提升32%,定位更精准且稀疏性增强35%-39%
- 适合需要可审计语音识别系统的研究与应用
基于Transformer的自动语音识别(ASR)模型如Whisper虽准确率高,但其预测难以解释。现有可解释AI(XAI)方法常缺乏忠实性与精确的时间对齐。本文提出一种模型内生的XAI框架LEAF-X,结合熵引导注意力加权、多层注意力传播及可选因果消融,识别低熵、高影响的注意力头与层,生成稀疏的词元到音帧归因。相比基于扰动的解释器或原始注意力图,LEAF-X利用编码器-解码器与语音增强型解码器仅模型的内部结构,生成更反映模型计算过程的解释。实验显示,其忠实度提升32%,局部性与稀疏性增强35%-39%,归因最稳定,支持更透明、可审计的ASR系统。
原文摘要 · Abstract (English)
Transformer-based automatic speech recognition (ASR) models such as Whisper are highly accurate, but their predictions remain difficult to interpret. Existing explainable AI (XAI) methods often lack faithfulness and precise temporal grounding. We propose Listening with Entropy-guided Attention for Faithful explainability (LEAF-X), a model-intrinsic XAI framework for transformer-based ASR. LEAF-X combines entropy-guided attention weighting, multi-layer attention rollout, and optional causal ablations to identify low-entropy, high-impact heads and layers, producing sparse token-to-frame attributions. Unlike perturbation-based explainers or raw attention maps, LEAF-X exploits the internal structure of encoder-decoder and speech-augmented decoder-only models to generate explanations that better reflect model computation. Results show 32% improved faithfulness, 35-39% stronger locality/sparsity, and the most stable attributions, supporting more transparent and auditable ASR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。