arXiv:2409.05015cs.HCcs.SD2024-09被引 10

通过声学适配与视觉对齐提升多模态情感识别效果。

Improving Multimodal Emotion Recognition by Leveraging Acoustic Adaptation and Visual Alignment

  • 用轻量微调优化语音模型各层,提升情感特征提取能力。
  • 利用无标签数据训练视觉编码器,使视觉特征对齐声学空间。
  • 融合声学、视觉与文本特征,测试集加权F1达88.90%。

多模态情感识别(MER)旨在通过融合多种模态信息自动识别和理解人类情感状态。然而,标注的多模态数据稀缺严重制约了该领域进展。本文针对MER 2024的MER-SEMI子挑战赛提出解决方案。首先,为更好适应声学模态特征用于情感识别,实验评估了预训练语音模型HuBERT不同层级在情感识别中的贡献;基于此,对最有效的层级进行参数高效微调(PEFT),以最少可学习参数实现最优情感识别适配。其次,利用声学模态优势,提出一种特征对齐预训练方法:使用大规模未标注数据训练视觉编码器,促进视觉特征在声学特征空间中的语义对齐。最后,结合适配后的声学特征、对齐的视觉特征与词法特征,采用注意力机制进行特征融合。在MER2024-SEMI测试集上,所提方法取得88.90%的加权F1分数,排名所有参赛团队第四,验证了方法有效性。

原文摘要 · Abstract (English)

Multimodal Emotion Recognition (MER) aims to automatically identify and understand human emotional states by integrating information from various modalities. However, the scarcity of annotated multimodal data significantly hinders the advancement of this research field. This paper presents our solution for the MER-SEMI sub-challenge of MER 2024. First, to better adapt acoustic modality features for the MER task, we experimentally evaluate the contributions of different layers of the pre-trained speech model HuBERT in emotion recognition. Based on these observations, we perform Parameter-Efficient Fine-Tuning (PEFT) on the layers identified as most effective for emotion recognition tasks, thereby achieving optimal adaptation for emotion recognition with a minimal number of learnable parameters. Second, leveraging the strengths of the acoustic modality, we propose a feature alignment pre-training method. This approach uses large-scale unlabeled data to train a visual encoder, thereby promoting the semantic alignment of visual features within the acoustic feature space. Finally, using the adapted acoustic features, aligned visual features, and lexical features, we employ an attention mechanism for feature fusion. On the MER2024-SEMI test set, the proposed method achieves a weighted F1 score of 88.90%, ranking fourth among all participating teams, validating the effectiveness of our approach.

情感识别多模态特征对齐语音建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。