通过对比学习与视觉序列压缩,提升多模态情感识别准确率。
A Novel Approach to for Multimodal Emotion Recognition : Multimodal semantic information fusion
- 结合对比学习与视觉序列压缩,优化跨模态特征融合
- 在IEMOCAP和MELD数据集上显著提升识别准确率
- 适合关注多模态融合与情感分析的研究者
随着人工智能与计算机视觉技术的发展,多模态情感识别成为研究热点。然而,现有方法在异构数据融合及模态相关性利用方面仍面临挑战。本文提出一种新方法DeepMSI-MER,融合对比学习与视觉序列压缩,通过对比学习增强跨模态特征融合,利用视觉序列压缩降低视觉模态冗余。在两个公开数据集IEMOCAP和MELD上的实验表明,DeepMSI-MER显著提升了情感识别的准确率与鲁棒性,验证了多模态特征融合及该方法的有效性。
原文摘要 · Abstract (English)
With the advancement of artificial intelligence and computer vision technologies, multimodal emotion recognition has become a prominent research topic. However, existing methods face challenges such as heterogeneous data fusion and the effective utilization of modality correlations. This paper proposes a novel multimodal emotion recognition approach, DeepMSI-MER, based on the integration of contrastive learning and visual sequence compression. The proposed method enhances cross-modal feature fusion through contrastive learning and reduces redundancy in the visual modality by leveraging visual sequence compression. Experimental results on two public datasets, IEMOCAP and MELD, demonstrate that DeepMSI-MER significantly improves the accuracy and robustness of emotion recognition, validating the effectiveness of multimodal feature fusion and the proposed approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。