通过跨模态注意力与对比学习提升多模态情感识别准确率。
MCN-CL: Multimodal Cross-Attention Network and Contrastive Learning for Multimodal Emotion Recognition

- 设计三查询机制与硬负样本挖掘,缓解模态差异与类别不平衡。
- 在IEMOCAP和MELD数据集上,加权F1分别提升3.42%和5.73%。
- 适合关注情感计算、跨模态融合的科研与工程人员。
多模态情感识别在心理健康监测、教育互动和人机交互等领域具有重要意义。然而,现有方法普遍面临三大挑战:类别分布不均、动态面部动作单元时间建模复杂,以及因模态异质性导致的特征融合困难。随着社交媒体中多模态数据的爆炸式增长,构建高效的跨模态融合框架以实现情感识别变得愈发迫切。为此,本文提出多模态交叉注意力网络与对比学习(MCN-CL)方法。该方法采用三查询机制与硬负样本挖掘策略,在去除特征冗余的同时保留关键情感线索,有效缓解了模态异质性与类别不平衡问题。在IEMOCAP和MELD数据集上的实验结果表明,所提方法优于当前最优方法,加权F1分数分别提升了3.42%和5.73%。
原文摘要 · Abstract (English)
Multimodal emotion recognition plays a key role in many domains, including mental health monitoring, educational interaction, and human-computer interaction. However, existing methods often face three major challenges: unbalanced category distribution, the complexity of dynamic facial action unit time modeling, and the difficulty of feature fusion due to modal heterogeneity. With the explosive growth of multimodal data in social media scenarios, the need for building an efficient cross-modal fusion framework for emotion recognition is becoming increasingly urgent. To this end, this paper proposes Multimodal Cross-Attention Network and Contrastive Learning (MCN-CL) for multimodal emotion recognition. It uses a triple query mechanism and hard negative mining strategy to remove feature redundancy while preserving important emotional cues, effectively addressing the issues of modal heterogeneity and category imbalance. Experiment results on the IEMOCAP and MELD datasets show that our proposed method outperforms state-of-the-art approaches, with Weighted F1 scores improving by 3.42% and 5.73%, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。