用高分辨率网络提升语音情感识别准确率
EmoHRNet: High-Resolution Neural Network Based Speech Emotion Recognition
- 采用保持高分辨率的HRNet结构提取语音特征
- 在RAVDESS等数据集上准确率达92.45%以上
- 适合需要精细情感分析的应用场景
语音情感识别(SER)对提升人机交互至关重要。本文提出EmoHRNet,一种针对SER优化的高分辨率网络(HRNet)新架构。该结构从初始层到最终层均保持高分辨率表示,将音频样本转换为频谱图后,利用HRNet提取高层次特征。其独特设计在整个过程中维持高分辨率表征,能够同时捕捉语音信号中的细微与整体情感线索。模型性能超越现有主流方法,在RAVDESS数据集上达到92.45%准确率,IEMOCAP上达80.06%,EMOVO上达92.77%。结果表明,EmoHRNet在SER领域树立了新基准。
原文摘要 · Abstract (English)
Speech emotion recognition (SER) is pivotal for enhancing human-machine interactions. This paper introduces "EmoHRNet", a novel adaptation of High-Resolution Networks (HRNet) tailored for SER. The HRNet structure is designed to maintain high-resolution representations from the initial to the final layers. By transforming audio samples into spectrograms, EmoHRNet leverages the HRNet architecture to extract high-level features. EmoHRNet's unique architecture maintains high-resolution representations throughout, capturing both granular and overarching emotional cues from speech signals. The model outperforms leading models, achieving accuracies of 92.45% on RAVDESS, 80.06% on IEMOCAP, and 92.77% on EMOVO. Thus, we show that EmoHRNet sets a new benchmark in the SER domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。