arXiv:2409.05007cs.SDcs.AI2024-09被引 4

用音频引导融合提升多模态情感分析,三甲方案。

Audio-Guided Fusion Techniques for Multimodal Emotion Analysis

  • 用音频模型Hubert-large引导跨通道与内通道信息融合
  • 通过自监督迭代优化,提升未标注数据利用效率
  • 针对数据分布不均设计先验投票机制,适合半监督场景

本文针对MER2024的半监督学习赛道(MER-SEMI)提出解决方案。首先,为增强情感分类性能,使用标注数据微调视频与文本特征提取器,即CLIP-vit-large和Baichuan-13B,有效保留视频中的原始情感信息。其次,提出一种音频引导的Transformer融合机制(AGT),利用Hubert-large的鲁棒性,显著提升跨模态与模态内信息融合效果。第三,通过高置信度无标签数据生成伪标签,迭代应用自监督学习以提高模型精度。最后,通过黑盒探针发现训练集与测试集存在数据分布不均问题,因此引入基于先验知识的投票机制。实验结果验证了该策略的有效性,最终在MER-SEMI赛道中获得第三名。

原文摘要 · Abstract (English)

In this paper, we propose a solution for the semi-supervised learning track (MER-SEMI) in MER2024. First, in order to enhance the performance of the feature extractor on sentiment classification tasks,we fine-tuned video and text feature extractors, specifically CLIP-vit-large and Baichuan-13B, using labeled data. This approach effectively preserves the original emotional information conveyed in the videos. Second, we propose an Audio-Guided Transformer (AGT) fusion mechanism, which leverages the robustness of Hubert-large, showing superior effectiveness in fusing both inter-channel and intra-channel information. Third, To enhance the accuracy of the model, we iteratively apply self-supervised learning by using high-confidence unlabeled data as pseudo-labels. Finally, through black-box probing, we discovered an imbalanced data distribution between the training and test sets. Therefore, We adopt a prior-knowledge-based voting mechanism. The results demonstrate the effectiveness of our strategy, ultimately earning us third place in the MER-SEMI track.

多模态情感分析半监督学习音频引导特征融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。