arXiv:2411.09189cs.AIcs.SD2024-11被引 4

双层LSTM提升语音情感识别准确率与实时性

Improvement and Implementation of a Speech Emotion Recognition Model Based on Dual-Layer LSTM

  • 采用双层LSTM捕捉语音序列长期依赖关系
  • 在RAVDESS数据集上准确率提升2%,延迟显著降低
  • 适合智能客服、人机交互等实时情感分析场景

本文在现有语音情感识别模型基础上,增加一层LSTM以提升音频数据的情感识别准确率与处理效率。通过双层LSTM网络捕捉语音序列中的长期依赖特征,模型能更准确识别复杂情绪模式。在RAVDESS数据集上的实验表明,该改进模型相比单层LSTM准确率提升2%,同时显著降低识别延迟,增强实时性能。结果表明,双层LSTM架构特别适用于处理具有长期依赖的情绪特征,为语音情感识别系统提供了有效优化方案,可应用于智能客服、情感分析和人机交互等领域。

原文摘要 · Abstract (English)

This paper builds upon an existing speech emotion recognition model by adding an additional LSTM layer to improve the accuracy and processing efficiency of emotion recognition from audio data. By capturing the long-term dependencies within audio sequences through a dual-layer LSTM network, the model can recognize and classify complex emotional patterns more accurately. Experiments conducted on the RAVDESS dataset validated this approach, showing that the modified dual layer LSTM model improves accuracy by 2% compared to the single-layer LSTM while significantly reducing recognition latency, thereby enhancing real-time performance. These results indicate that the dual-layer LSTM architecture is highly suitable for handling emotional features with long-term dependencies, providing a viable optimization for speech emotion recognition systems. This research provides a reference for practical applications in fields like intelligent customer service, sentiment analysis and human-computer interaction.

语音识别情感分析LSTM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。