用中心损失提升语音情绪识别的特征区分度
learning discriminative features from spectrograms using center loss for speech emotion recognition
- 联合交叉熵与中心损失,让不同情绪特征更易分离
- 在梅尔频谱上准确率提升超3%,短时傅里叶频谱超4%
- 适合做语音情绪识别特征优化的研究者参考
从语音中识别情绪对实现人机自然交互至关重要。然而,由于情绪本身具有模糊性,提取有效特征仍具挑战。本文提出一种新方法,通过协同使用 softmax 交叉熵损失与中心损失,从长度可变的频谱图中学习更具区分性的特征。交叉熵损失使不同情绪类别的特征相互分离,中心损失则将同类别特征拉向其类别中心。两者结合显著增强特征判别力,使网络学习到更有效的情绪识别特征。实验表明,引入中心损失后,在梅尔频谱输入下,未加权准确率和加权准确率均提升超过3%;在短时傅里叶变换频谱输入下,提升超过4%。
原文摘要 · Abstract (English)
Identifying the emotional state from speech is essential for the natural interaction of the machine with the speaker. However, extracting effective features for emotion recognition is difficult, as emotions are ambiguous. We propose a novel approach to learn discriminative features from variable length spectrograms for emotion recognition by cooperating softmax cross-entropy loss and center loss together. The softmax cross-entropy loss enables features from different emotion categories separable, and center loss efficiently pulls the features belonging to the same emotion category to their center. By combining the two losses together, the discriminative power will be highly enhanced, which leads to network learning more effective features for emotion recognition. As demonstrated by the experimental results, after introducing center loss, both the unweighted accuracy and weighted accuracy are improved by over 3\% on Mel-spectrogram input, and more than 4\% on Short Time Fourier Transform spectrogram input.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。