通过特征分布对齐提升多模态语音情绪识别准确率
Multi-modal Speech Emotion Recognition via Feature Distribution Adaptation Network
- 用跨模态注意力建模音视频特征相似性
- 在两个基准数据集上优于现有方法
- 适合做跨模态情感分析的研究者参考
本文提出一种新型深度归纳迁移学习框架——特征分布适配网络,用于解决多模态语音情绪识别难题。该方法利用深度迁移学习策略对齐视觉与音频特征分布,获得一致的情绪表征,从而提升识别性能。模型中分别使用预训练的ResNet-34提取面部表情图像和声学梅尔频谱图的特征;引入交叉注意力机制建模多模态特征间的内在相似关系;最后通过前馈网络结合局部最大均值差异损失,高效实现多模态特征分布适配。在两个基准数据集上的实验表明,所提模型相较现有方法表现优异。
原文摘要 · Abstract (English)
In this paper, we propose a novel deep inductive transfer learning framework, named feature distribution adaptation network, to tackle the challenging multi-modal speech emotion recognition problem. Our method aims to use deep transfer learning strategies to align visual and audio feature distributions to obtain consistent representation of emotion, thereby improving the performance of speech emotion recognition. In our model, the pre-trained ResNet-34 is utilized for feature extraction for facial expression images and acoustic Mel spectrograms, respectively. Then, the cross-attention mechanism is introduced to model the intrinsic similarity relationships of multi-modal features. Finally, the multi-modal feature distribution adaptation is performed efficiently with feed-forward network, which is extended using the local maximum mean discrepancy loss. Experiments are carried out on two benchmark datasets, and the results demonstrate that our model can achieve excellent performance compared with existing ones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。