arXiv:2506.02181cs.CLcs.AI2025-06中稿 · Interspeech 2025被引 2

用归因分析揭示现代语音识别模型依赖的声学线索。

Echoes of Phonetics: Unveiling Relevant Acoustic Cues for ASR via Feature Attribution

  • 通过特征归因技术解析模型关注的声学特征
  • 发现模型更依赖元音完整时长和前两个共振峰
  • 适合想提升语音识别可解释性的研究者

尽管语音识别(ASR)取得显著进展,但模型依赖的具体声学线索仍不明确。以往研究仅针对少数音素和过时模型展开。本文采用特征归因方法,分析现代基于Conformer的ASR系统对塞音、擦音和元音的依赖。结果表明,该模型依赖元音的完整时间跨度,尤其是前两个共振峰,在男声中更具显著性;对嘶音擦音的频谱特征捕捉优于非嘶音擦音;在塞音中更关注释放阶段,尤其注重爆发特征。这些发现提升了ASR模型的可解释性,并指出了未来提升模型鲁棒性的研究方向。

原文摘要 · Abstract (English)

Despite significant advances in ASR, the specific acoustic cues models rely on remain unclear. Prior studies have examined such cues on a limited set of phonemes and outdated models. In this work, we apply a feature attribution technique to identify the relevant acoustic cues for a modern Conformer-based ASR system. By analyzing plosives, fricatives, and vowels, we assess how feature attributions align with their acoustic properties in the time and frequency domains, also essential for human speech perception. Our findings show that the ASR model relies on vowels' full time spans, particularly their first two formants, with greater saliency in male speech. It also better captures the spectral characteristics of sibilant fricatives than non-sibilants and prioritizes the release phase in plosives, especially burst characteristics. These insights enhance the interpretability of ASR models and highlight areas for future research to uncover potential gaps in model robustness.

语音识别可解释性声学特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。