arXiv:2504.05368cs.SDeess.AS2025-04被引 7

提出首个用于语音情感识别的局部可解释方法,解析关键频段贡献。

Exploring Local Interpretable Model-Agnostic Explanations for Speech Emotion Recognition with Distribution-Shift

  • 基于LIME框架,生成语音情感模型的可解释频段分析。
  • 在三个数据集上验证,对模型差异更鲁棒,对分布偏移敏感。
  • 适合需要理解模型决策依据的语音情感研究者使用。

我们提出了EmoLIME,一种针对黑箱语音情感识别(SER)模型的局部可解释模型无关解释方法。据我们所知,这是首次将LIME应用于SER任务。EmoLIME生成高层次可解释性解释,识别影响情感判断的具体频率范围,有助于解读端到端语音模型产生的高维嵌入。我们在三个情感语音数据集上,对基于手工声学特征和Wav2Vec 2.0嵌入训练的分类器,进行了定性、定量和统计评估。结果表明,EmoLIME在不同模型间表现出更强的鲁棒性,而在不同数据集分布偏移下表现较弱,凸显其在单数据集内实现一致解释的潜力。

原文摘要 · Abstract (English)

We introduce EmoLIME, a version of local interpretable model-agnostic explanations (LIME) for black-box Speech Emotion Recognition (SER) models. To the best of our knowledge, this is the first attempt to apply LIME in SER. EmoLIME generates high-level interpretable explanations and identifies which specific frequency ranges are most influential in determining emotional states. The approach aids in interpreting complex, high-dimensional embeddings such as those generated by end-to-end speech models. We evaluate EmoLIME, qualitatively, quantitatively, and statistically, across three emotional speech datasets, using classifiers trained on both hand-crafted acoustic features and Wav2Vec 2.0 embeddings. We find that EmoLIME exhibits stronger robustness across different models than across datasets with distribution shifts, highlighting its potential for more consistent explanations in SER tasks within a dataset.

语音情感可解释性LIME频段分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。