arXiv:2507.07046cs.SDcs.AI2025-07被引 2

用混合模型提升语音情绪识别准确率,跨数据集表现优异。

A Novel Hybrid Deep Learning Technique for Speech Emotion Detection using Feature Engineering

  • 结合特征工程与双向LSTM构建DCRF-BiLSTM模型
  • 在五数据集组合上达93.76%准确率,单个数据集最高100%
  • 首次统一评估所有基准数据集,适合多场景情绪识别研究

当前语音情绪识别(SER)在人机交互(HCI)和人工智能(AI)发展中至关重要。本文提出的DCRF-BiLSTM模型用于识别七类情绪:中性、快乐、悲伤、愤怒、恐惧、厌恶和惊讶,训练于RAVDESS(R)、TESS(T)、SAVEE(S)、EmoDB(E)和Crema-D(C)五个数据集。模型在各数据集上均表现优异,包括在RAVDESS上达到97.83%,SAVEE上97.02%,CREMA-D上95.10%,TESS和EmoDB上均为100%。在(R+T+S)组合数据集上达到98.82%准确率,优于已有结果。据我们所知,尚未有研究同时在全部五个基准数据集(即R+T+S+C+E)上评估单一SER模型。本文首次引入该全面组合,实现93.76%的整体准确率。这些结果证实了DCRF-BiLSTM框架在多样化数据集上的鲁棒性与泛化能力。

原文摘要 · Abstract (English)

Nowadays, speech emotion recognition (SER) plays a vital role in the field of human-computer interaction (HCI) and the evolution of artificial intelligence (AI). Our proposed DCRF-BiLSTM model is used to recognize seven emotions: neutral, happy, sad, angry, fear, disgust, and surprise, which are trained on five datasets: RAVDESS (R), TESS (T), SAVEE (S), EmoDB (E), and Crema-D (C). The model achieves high accuracy on individual datasets, including 97.83% on RAVDESS, 97.02% on SAVEE, 95.10% for CREMA-D, and a perfect 100% on both TESS and EMO-DB. For the combined (R+T+S) datasets, it achieves 98.82% accuracy, outperforming previously reported results. To our knowledge, no existing study has evaluated a single SER model across all five benchmark datasets (i.e., R+T+S+C+E) simultaneously. In our work, we introduce this comprehensive combination and achieve a remarkable overall accuracy of 93.76%. These results confirm the robustness and generalizability of our DCRF-BiLSTM framework across diverse datasets.

语音识别情绪检测深度学习特征工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。