交叉注意力仅解释一半语音到文本模型的决策依据。
Cross-Attention is Half Explanation in Speech-to-Text Models
- 用特征归因法对比注意力得分,评估其解释力。
- 注意力仅捕捉约50%输入相关性,最高覆盖75%相关性。
- 适合关注模型可解释性与注意力机制局限的研究者。
交叉注意力是编码器-解码器架构中的核心机制,广泛应用于语音到文本(S2T)处理等领域。其得分常被用于时间戳估计、音文对齐等下游任务,假设其反映了输入语音表示与生成文本间的依赖关系。尽管注意力机制的解释性在自然语言处理领域备受争议,但在语音领域仍缺乏系统研究。本文通过对比交叉注意力得分与基于特征归因得到的输入显著性图,评估其解释能力。分析涵盖单语、多语、单任务与多任务模型,覆盖多个规模层级。结果表明,注意力得分在跨头和层聚合后与显著性解释中度至强相关,但仅能捕捉约50%的输入相关性,在最佳情况下仅反映52%-75%的解码器注意力机制。这揭示了将交叉注意力视为解释代理的根本局限,表明其仅提供信息丰富但不完整的预测驱动因素视图。
原文摘要 · Abstract (English)
Cross-attention is a core mechanism in encoder-decoder architectures, widespread in many fields, including speech-to-text (S2T) processing. Its scores have been repurposed for various downstream applications--such as timestamp estimation and audio-text alignment--under the assumption that they reflect the dependencies between input speech representation and the generated text. While the explanatory nature of attention mechanisms has been widely debated in the broader NLP literature, this assumption remains largely unexplored within the speech domain. To address this gap, we assess the explanatory power of cross-attention in S2T models by comparing its scores to input saliency maps derived from feature attribution. Our analysis spans monolingual and multilingual, single-task and multi-task models at multiple scales, and shows that attention scores moderately to strongly align with saliency-based explanations, particularly when aggregated across heads and layers. However, it also shows that cross-attention captures only about 50% of the input relevance and, in the best case, only partially reflects how the decoder attends to the encoder's representations--accounting for just 52-75% of the saliency. These findings uncover fundamental limitations in interpreting cross-attention as an explanatory proxy, suggesting that it offers an informative yet incomplete view of the factors driving predictions in S2T models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。