arXiv:2409.18971cs.MMcs.AI2024-09被引 8

早期融合音文特征,提升多模态情绪识别准确率

Early Joint Learning of Emotion Information Makes MultiModal Model Understand You Better

  • 用大模型实现音文早期联合训练,缓解模态竞争
  • 在两个子挑战中均获第二,验证方法有效性
  • 适合关注多模态情感分析的研究者和开发者

本文针对多模态情绪识别挑战赛(MER2024)的两个子任务提出解决方案。为缓解音频与文本之间的模态竞争问题,采用基于大语言模型的早期融合策略,先对音频和文本进行联合训练,再将联合特征与单模态特征晚期融合。为应对数据不足和类别不平衡问题,采用多轮多模型投票进行数据挖掘。同时,通过语音源分离预处理音频,提升音频特征质量。所提模型在MER2024-SEMI和MER2024-NOISE两个子任务中均位列第2,验证了方法的有效性与鲁棒性。

原文摘要 · Abstract (English)

In this paper, we present our solutions for emotion recognition in the sub-challenges of Multimodal Emotion Recognition Challenge (MER2024). To mitigate the modal competition issue between audio and text, we adopt an early fusion strategy based on a large language model, where joint training of audio and text is conducted initially. And the joint Audio-Text modal feature will be late-fused with other unimodal features. In order to solve the problems of data insufficiency and class imbalance, We use multiple turns of multi-model voting for data mining. Moreover, to enhance the quality of audio features, we employ speech source separation to preprocess audios. Our model ranks \textbf{2nd} in both MER2024-SEMI and MER2024-NOISE, validating our method's effectiveness.

多模态情绪识别语音处理早期融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。