arXiv:2509.26543cs.CLcs.AI2025-09中稿 · BlackBoxNLP 2025被引 3

首次为语音转文本模型提供对比解释,揭示音频特征如何影响输出选择。

The Unheard Alternative: Contrastive Explanations for Speech-to-Text Models

  • 基于输入频谱图的特征归因,分析不同输出间的差异驱动因素。
  • 在语音翻译性别指代案例中准确识别出影响性别选择的关键音频特征。
  • 适合关注语音模型可解释性与公平性的研究者和开发者。

对比解释通过说明AI系统为何产生某一输出而非另一输出,被广泛认为比传统解释更具信息量和可读性。然而,获取语音转文本(S2T)生成模型的对比解释仍是一个开放挑战。本文借鉴特征归因技术,提出首个适用于S2T的对比解释方法,通过分析输入频谱图的不同部分如何影响替代输出的选择。在性别指代的语音翻译案例研究中,该方法能准确识别出决定性别选择的关键音频特征。本工作将对比解释的适用范围拓展至S2T领域,为深入理解此类模型提供了基础。

原文摘要 · Abstract (English)

Contrastive explanations, which indicate why an AI system produced one output (the target) instead of another (the foil), are widely regarded in explainable AI as more informative and interpretable than standard explanations. However, obtaining such explanations for speech-to-text (S2T) generative models remains an open challenge. Drawing from feature attribution techniques, we propose the first method to obtain contrastive explanations in S2T by analyzing how parts of the input spectrogram influence the choice between alternative outputs. Through a case study on gender assignment in speech translation, we show that our method accurately identifies the audio features that drive the selection of one gender over another. By extending the scope of contrastive explanations to S2T, our work provides a foundation for better understanding S2T models.

可解释AI语音转文本对比解释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。