arXiv:2507.03251cs.SDcs.AI2025-07被引 1

用梅尔倒谱系数+注意力机制提升语音情感识别精度

Toward Efficient Speech Emotion Recognition via Spectral Learning and Attention

  • 以梅尔倒谱系数为特征,结合1D-CNN与通道/空间注意力
  • 在6个数据集上最高达99.82%准确率,刷新基准
  • 适合做智能客服、辅助医疗等实时情感交互场景

语音情感识别(SER)传统依赖听觉数据分析情绪分类。现有方法常难以捕捉细微情感差异且泛化能力不足。本文采用梅尔倒谱系数(MFCCs)作为频谱特征,连接计算情感处理与人类听觉感知。提出一种基于1D-CNN的SER框架,融合数据增强技术。从增强数据中提取的MFCC特征经改进的1D-CNN处理,该网络引入通道与空间注意力机制,使模型聚焦关键情感模式,增强对语音信号细微变化的捕捉能力。所提方法在多个公开数据集上表现卓越:SAVEE达到97.49%,RAVDESS为99.23%,CREMA-D为89.31%,TESS为99.82%,EMO-DB为99.53%,EMOVO为96.39%。实验结果表明,该方法显著提升跨数据集泛化能力,为助人技术与人机交互中的实际部署提供高精度解决方案。

原文摘要 · Abstract (English)

Speech Emotion Recognition (SER) traditionally relies on auditory data analysis for emotion classification. Several studies have adopted different methods for SER. However, existing SER methods often struggle to capture subtle emotional variations and generalize across diverse datasets. In this article, we use Mel-Frequency Cepstral Coefficients (MFCCs) as spectral features to bridge the gap between computational emotion processing and human auditory perception. To further improve robustness and feature diversity, we propose a novel 1D-CNN-based SER framework that integrates data augmentation techniques. MFCC features extracted from the augmented data are processed using a 1D Convolutional Neural Network (CNN) architecture enhanced with channel and spatial attention mechanisms. These attention modules allow the model to highlight key emotional patterns, enhancing its ability to capture subtle variations in speech signals. The proposed method delivers cutting-edge performance, achieving the accuracy of 97.49% for SAVEE, 99.23% for RAVDESS, 89.31% for CREMA-D, 99.82% for TESS, 99.53% for EMO-DB, and 96.39% for EMOVO. Experimental results show new benchmarks in SER, demonstrating the effectiveness of our approach in recognizing emotional expressions with high precision. Our evaluation demonstrates that the integration of advanced Deep Learning (DL) methods substantially enhances generalization across diverse datasets, underscoring their potential to advance SER for real-world deployment in assistive technologies and human-computer interaction.

语音情感识别梅尔倒谱系数注意力机制深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。