arXiv:2409.07078cs.CVcs.AI2024-09被引 16

用视觉语言提示和模态丢弃提升视频情感识别准确率

Multimodal Emotion Recognition with Vision-language Prompting and Modality Dropout

  • 基于CLIP构建情绪感知的EmoVCLIP,结合视觉-语言提示学习
  • 通过模态丢弃增强多模态融合鲁棒性,测试集准确率达90.15%
  • 适合关注视频情感分析与多模态模型优化的研究者

本文针对第二届多模态情感识别挑战赛赛道1(MER2024-SEMI)提出解决方案。为提升情感识别的准确率与泛化能力,我们提出多项方法:首先,引入EmoVCLIP,该模型基于CLIP,采用视觉-语言提示学习进行微调,专用于视频情感识别任务;通过在预训练CLIP上应用提示学习,显著提升了其在情感视频上的表现。其次,为缓解多模态融合中的模态依赖问题,采用模态丢弃策略以实现更鲁棒的信息融合。此外,为辅助Baichuan更好地提取情绪信息,建议使用GPT-4生成提示文本作为输入。最后,采用自训练策略,利用模型生成的高置信度伪标签,将未标注视频加入训练集。实验结果表明,本模型在MER2024-SEMI赛道排名第一,测试集准确率达到90.15%。

原文摘要 · Abstract (English)

In this paper, we present our solution for the Second Multimodal Emotion Recognition Challenge Track 1(MER2024-SEMI). To enhance the accuracy and generalization performance of emotion recognition, we propose several methods for Multimodal Emotion Recognition. Firstly, we introduce EmoVCLIP, a model fine-tuned based on CLIP using vision-language prompt learning, designed for video-based emotion recognition tasks. By leveraging prompt learning on CLIP, EmoVCLIP improves the performance of pre-trained CLIP on emotional videos. Additionally, to address the issue of modality dependence in multimodal fusion, we employ modality dropout for robust information fusion. Furthermore, to aid Baichuan in better extracting emotional information, we suggest using GPT-4 as the prompt for Baichuan. Lastly, we utilize a self-training strategy to leverage unlabeled videos. In this process, we use unlabeled videos with high-confidence pseudo-labels generated by our model and incorporate them into the training set. Experimental results demonstrate that our model ranks 1st in the MER2024-SEMI track, achieving an accuracy of 90.15% on the test set.

情感识别多模态提示学习自训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。