arXiv:2409.04447cs.SDcs.AI2024-09中稿 · ACM MM Workshop 20…被引 15

用对比学习与自训练,在少量标注数据下提升多模态情绪识别效果。

Leveraging Contrastive Learning and Self-Training for Multimodal Emotion Recognition with Limited Labeled Samples

  • 基于三模态数据设计组合对比学习框架,增强模型表征能力。
  • 在有限标注数据下实现88.25%的加权F1分数,排名第六。
  • 适合数据稀缺场景下的多模态情绪识别研究者参考。

多模态情绪识别挑战MER2024聚焦于利用语音、语言和视觉信号识别情绪。本文提出针对半监督学习子挑战(MER2024-SEMI)的解决方案,应对情绪识别中标注数据不足的问题。首先,为缓解类别不平衡,采用过采样策略;其次,提出一种在三模态输入数据上构建的模态表示组合对比学习(MR-CCL)框架,以建立鲁棒的初始模型;第三,探索自训练方法扩充训练集;最后,通过多分类器加权软投票策略提升预测鲁棒性。所提方法在MER2024-SEMI挑战中验证有效,取得88.25%的加权平均F-score, leaderboard 排名第6。项目代码见 https://github.com/WooyoohL/MER2024-SEMI。

原文摘要 · Abstract (English)

The Multimodal Emotion Recognition challenge MER2024 focuses on recognizing emotions using audio, language, and visual signals. In this paper, we present our submission solutions for the Semi-Supervised Learning Sub-Challenge (MER2024-SEMI), which tackles the issue of limited annotated data in emotion recognition. Firstly, to address the class imbalance, we adopt an oversampling strategy. Secondly, we propose a modality representation combinatorial contrastive learning (MR-CCL) framework on the trimodal input data to establish robust initial models. Thirdly, we explore a self-training approach to expand the training set. Finally, we enhance prediction robustness through a multi-classifier weighted soft voting strategy. Our proposed method is validated to be effective on the MER2024-SEMI Challenge, achieving a weighted average F-score of 88.25% and ranking 6th on the leaderboard. Our project is available at https://github.com/WooyoohL/MER2024-SEMI.

情绪识别对比学习自训练多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。