arXiv:2506.00903cs.CV2025-06中稿 · IEEE/CVF WACV 2025…被引 14

用CLIP的语义知识提升多模态情感识别效果

Leveraging CLIP Encoder for Multimodal Emotion Recognition

  • 用标签文本嵌入引导跨模态特征学习
  • 在CMU-MOSI和CMU-MOSEI上超越现有方法
  • 适合需要小样本情感识别的研究者

多模态情感识别(MER)旨在通过语言、音频和视觉等多源数据识别人类情绪。尽管近年进展显著,但大规模标注数据的缺乏仍限制性能提升。为此,本文基于对比语言-图像预训练(CLIP)模型及其在海量数据中学习到的语义知识,提出一种标签编码器引导的多模态情感识别框架(MER-CLIP)。该方法将标签视为文本嵌入,以融合其语义信息,从而学习更具代表性的情感特征。进一步设计跨模态解码器,通过标签编码器提供的语义引导,逐步融合各模态特征并映射至共享嵌入空间。最终,标签编码器引导的预测机制通过嵌入标签语义实现对多样标签的泛化能力。实验表明,该方法在基准数据集CMU-MOSI和CMU-MOSEI上优于当前最优方法。

原文摘要 · Abstract (English)

Multimodal emotion recognition (MER) aims to identify human emotions by combining data from various modalities such as language, audio, and vision. Despite the recent advances of MER approaches, the limitations in obtaining extensive datasets impede the improvement of performance. To mitigate this issue, we leverage a Contrastive Language-Image Pre-training (CLIP)-based architecture and its semantic knowledge from massive datasets that aims to enhance the discriminative multimodal representation. We propose a label encoder-guided MER framework based on CLIP (MER-CLIP) to learn emotion-related representations across modalities. Our approach introduces a label encoder that treats labels as text embeddings to incorporate their semantic information, leading to the learning of more representative emotional features. To further exploit label semantics, we devise a cross-modal decoder that aligns each modality to a shared embedding space by sequentially fusing modality features based on emotion-related input from the label encoder. Finally, the label encoder-guided prediction enables generalization across diverse labels by embedding their semantic information as well as word labels. Experimental results show that our method outperforms the state-of-the-art MER methods on the benchmark datasets, CMU-MOSI and CMU-MOSEI.

多模态情感识别CLIP语义引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。