通过多路径跨模态交互,提升多模态情感识别准确率。
MCIHN: A Hybrid Network Model Based on Multi-path Cross-modal Interaction for Multimodal Emotion Recognition
- 分模态构建对抗自编码器提取情绪特征
- 跨模态门控机制减少模态差异,生成交互特征
- 在SIMS和MOSI数据集上表现优于现有方法
多模态情感识别对未来的智能人机交互至关重要。然而,由于各模态间差异以及单模态情感信息难以刻画,准确识别仍面临挑战。为此,提出一种基于多路径跨模态交互的混合网络模型(MCIHN)。首先,为每个模态分别构建对抗自编码器(AAE),通过编码器学习判别性情绪特征,并利用解码器重构特征以获得更丰富的类别判别信息。随后,将各模态的潜在编码输入预定义的跨模态门控机制(CGMM),降低模态间差异,建立模态间的语义关联,并生成跨模态交互特征。最后,通过特征融合模块(FFM)进行多模态融合,提升情感识别性能。在公开的SIMS和MOSI数据集上的实验表明,MCIHN取得更优效果。
原文摘要 · Abstract (English)
Multimodal emotion recognition is crucial for future human-computer interaction. However, accurate emotion recognition still faces significant challenges due to differences between different modalities and the difficulty of characterizing unimodal emotional information. To solve these problems, a hybrid network model based on multipath cross-modal interaction (MCIHN) is proposed. First, adversarial autoencoders (AAE) are constructed separately for each modality. The AAE learns discriminative emotion features and reconstructs the features through a decoder to obtain more discriminative information about the emotion classes. Then, the latent codes from the AAE of different modalities are fed into a predefined Cross-modal Gate Mechanism model (CGMM) to reduce the discrepancy between modalities, establish the emotional relationship between interacting modalities, and generate the interaction features between different modalities. Multimodal fusion using the Feature Fusion module (FFM) for better emotion recognition. Experiments were conducted on publicly available SIMS and MOSI datasets, demonstrating that MCIHN achieves superior performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。