轻量级模型提升语音情感识别准确率,兼顾效率与细节捕捉
Enhanced Speech Emotion Recognition with Efficient Channel Attention Guided Deep CNN-BiLSTM Framework
- 用注意力机制提取语音局部与全局特征
- 跨5个数据集平均准确率达97.19%~99.65%
- 适合需要高效实时情感分析的交互系统
语音情感识别(SER)对情感计算和人机交互至关重要。现有方法在低计算成本下难以有效提取关键特征。本文提出一种轻量级SER架构,融合注意力引导的局部特征块(ALFB)捕捉语音信号中的高级相关特征,并引入全局特征块(GFB)捕获序列性与长时依赖信息。通过聚合局部与全局注意力特征向量,模型有效建模了反映复杂情感线索的显著特征内在关联。实验中提取梅尔频率倒谱系数、梅尔频谱图、均方根值和过零率四类谱特征,采用五折交叉验证,在TESS、RAVDESS、BanglaSER、SUBESCO和Emo-DB五个多语言标准数据集上分别取得99.65%、94.88%、98.12%、97.94%和97.19%的平均准确率,性能优于多数现有方法。
原文摘要 · Abstract (English)
Speech emotion recognition (SER) is crucial for enhancing affective computing and enriching the domain of human-computer interaction. However, the main challenge in SER lies in selecting relevant feature representations from speech signals with lower computational costs. In this paper, we propose a lightweight SER architecture that integrates attention-based local feature blocks (ALFBs) to capture high-level relevant feature vectors from speech signals. We also incorporate a global feature block (GFB) technique to capture sequential, global information and long-term dependencies in speech signals. By aggregating attention-based local and global contextual feature vectors, our model effectively captures the internal correlation between salient features that reflect complex human emotional cues. To evaluate our approach, we extracted four types of spectral features from speech audio samples: mel-frequency cepstral coefficients, mel-spectrogram, root mean square value, and zero-crossing rate. Through a 5-fold cross-validation strategy, we tested the proposed method on five multi-lingual standard benchmark datasets: TESS, RAVDESS, BanglaSER, SUBESCO, and Emo-DB, and obtained a mean accuracy of 99.65%, 94.88%, 98.12%, 97.94%, and 97.19% respectively. The results indicate that our model achieves state-of-the-art (SOTA) performance compared to most existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。