arXiv:2607.16803cs.SDcs.AI2026-07中稿 · ICML

轻量级语音情感识别模型,兼顾准确率与可解释性。

Explainable Lightweight Compact Deep Models for Speech Emotion Recognition

论文配图:Explainable Lightweight Compact Deep Models for Speech Emotion Recognition
图 1 · 摘自论文原文
  • 用紧凑卷积网络+注意力统计池化捕捉情绪关键时频特征。
  • 在SAVEE数据集上参数量显著减少,准确率仍具竞争力。
  • 结合Grad-CAM可视化预测依据,适合医疗等透明性要求高场景。

语音情感识别(SER)在医疗、客服及人机交互等以用户为中心的应用中至关重要。在医疗与辅助决策场景中,人们越来越关注既能准确识别情绪,又能提供透明预测且易于部署的模型。然而,现有大多数SER方法依赖复杂的深度学习架构,导致可解释性差且计算开销大。本文提出一种基于紧凑卷积神经网络的可解释轻量级语音情感识别框架。该方法采用对数梅尔频谱图表示以捕捉语音的时频特征,并利用注意力统计池化突出情绪显著的时间段。为提升模型透明度,引入基于梯度的类激活映射(Grad-CAM)来可视化影响预测的关键时间-频率区域。在SAVEE情感语音数据集上的实验表明,所提框架在保持紧凑结构的同时实现了具有竞争力的识别性能,参数量远低于多数现有SER模型。结果表明,高效卷积架构结合可解释分析,可在识别准确率、计算效率与模型透明性之间实现良好平衡。

原文摘要 · Abstract (English)

Speech Emotion Recognition (SER) is an important component in a wide range of human-centered applications, including healthcare, customer service, and human-omputer interaction. In medical and decision-support settings, there is increasing interest in models that not only achieve accurate emotion recognition but also support transparent predictions and efficient deployment. However, many existing SER approaches rely on complex deep learning architectures that limit interpretability and increase computational cost. This paper presents an explainable and lightweight speech emotion recognition framework based on a compact convolutional neural network architecture. The proposed approach utilizes log-Mel spectrogram representations to capture spectro-temporal speech characteristics and employs attentive statistics pooling to emphasize emotionally salient temporal segments. To improve model transparency, gradient-based class activation mapping (Grad-CAM) is incorporated to visualize the time-frequency regions that influence the model's predictions. Experimental evaluation on the SAVEE emotional speech dataset demonstrates that the proposed framework achieves competitive recognition performance while maintaining a compact architecture with significantly fewer parameters than many existing SER models. The results indicate that efficient convolutional architectures combined with interpretable analysis can provide a practical balance between recognition accuracy, computational efficiency, and model transparency.

语音识别轻量模型可解释性情感分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。