针对中文情感与意图联合识别中的类别不平衡问题,提出多模态增强与加权焦点对比损失方法。
Mitigating Category Imbalance: Fosafer System for the Multimodal Emotion and Intent Joint Understanding Challenge
- 融合文本、视频、音频的多模态数据增强,缓解类别不平衡。
- 采用样本加权焦点对比损失,提升少数类和难区分样本的识别效果。
- 适合关注多模态情感分析、跨模态学习的研究者或应用开发者。
本文针对中文多模态情感与意图联合理解挑战赛的第二赛道,提出Fosafer方法,旨在解决类别不平衡问题下的情感与意图联合识别。为缓解该问题,我们在文本、视频和音频模态上应用多种数据增强技术。同时,提出样本加权焦点对比损失(SampleWeighted Focal Contrastive loss),以应对少数类样本及语义相近但难以区分样本的识别挑战。此外,对Hubert模型进行微调以适应情感与意图联合识别任务。为减少模态间竞争,引入模态丢弃策略。最终预测采用多数投票机制。实验结果表明,本方法在第二赛道中取得第二名成绩。
原文摘要 · Abstract (English)
This paper presents Fosafer approach to the Track 2 Mandarin in the Multimodal Emotion and Intent Joint Understandingchallenge, which focuses on achieving joint recognition of emotion and intent in Mandarin, despite the issue of category imbalance. To alleviate this issue, we use a variety of data augmentation techniques across text, video, and audio modalities. Additionally, we introduce the SampleWeighted Focal Contrastive loss, designed to address the challenges of recognizing minority class samples and those that are semantically similar but difficult to distinguish. Moreover, we fine-tune the Hubert model to adapt the emotion and intent joint recognition. To mitigate modal competition, we introduce a modal dropout strategy. For the final predictions, a plurality voting approach is used to determine the results. The experimental results demonstrate the effectiveness of our method, which achieves the second-best performance in the Track 2 Mandarin challenge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。