用专家声学线索增强语音情绪识别解释,让模型更可信
Beyond saliency: enhancing explanation of speech emotion recognition with expert-referenced acoustic cues
- 将显著区域与专家定义的声学特征关联,明确为何重要
- 在基准数据集上提升解释质量,使结果更可信
- 适合关注可解释性与情绪计算可信度的研究者
语音情绪识别(SER)的可解释人工智能(XAI)对构建透明、可信的模型至关重要。现有基于显著性的方法源自视觉领域,仅标注频谱图区域,却无法验证这些区域是否对应真实的情绪声学标志,限制了其忠实度与可解释性。本文提出一种新框架,通过量化显著区域内声学线索的强度,明确‘被强调的是什么’并连接‘为什么重要’,将显著性与专家参考的语音情绪声学特征相联系。在多个基准SER数据集上的实验表明,该方法显著提升了解释质量,使模型输出更易理解且符合人类专家认知。相比传统显著性方法,本方案提供了更合理、更可信的语音情绪识别解释,为可信赖的情绪计算奠定基础。
原文摘要 · Abstract (English)
Explainable AI (XAI) for Speech Emotion Recognition (SER) is critical for building transparent, trustworthy models. Current saliency-based methods, adapted from vision, highlight spectrogram regions but fail to show whether these regions correspond to meaningful acoustic markers of emotion, limiting faithfulness and interpretability. We propose a framework that overcomes these limitations by quantifying the magnitudes of cues within salient regions. This clarifies "what" is highlighted and connects it to "why" it matters, linking saliency to expert-referenced acoustic cues of speech emotions. Experiments on benchmark SER datasets show that our approach improves explanation quality by explicitly linking salient regions to theory-driven speech emotions expert-referenced acoustics. Compared to standard saliency methods, it provides more understandable and plausible explanations of SER models, offering a foundational step towards trustworthy speech-based affective computing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。