通过多模态融合提升情感识别准确率,解决数据不足问题。
ECMF: Enhanced Cross-Modal Fusion for Multimodal Emotion Recognition in MER-SEMI Challenge
- 设计双分支视觉编码器与上下文增强文本方法,提取多模态特征。
- 融合策略实现动态加权与原始特征保留,提升识别效果。
- 适用于人机交互中的情感分析,尤其适合小样本场景。
情感识别在提升人机交互中至关重要。本文针对MER2025竞赛中的MER-SEMI挑战,提出一种新型多模态情感识别框架。为应对数据稀缺问题,利用大规模预训练模型从视觉、音频和文本模态中提取有效特征。视觉模态采用双分支编码器,捕捉全局帧级特征与局部面部表征;文本模态引入基于大语言模型的上下文增强方法,丰富输入文本中的情感线索。为有效融合多模态特征,提出包含自注意力机制(实现动态模态加权)和残差连接(保留原始表示)的融合策略。此外,通过多源标签策略优化训练集中的噪声标签。该方法在MER2025-SEMI数据集上相较官方基线显著提升性能,加权F-score达87.49%,高于基线78.63%,验证了所提框架的有效性。
原文摘要 · Abstract (English)
Emotion recognition plays a vital role in enhancing human-computer interaction. In this study, we tackle the MER-SEMI challenge of the MER2025 competition by proposing a novel multimodal emotion recognition framework. To address the issue of data scarcity, we leverage large-scale pre-trained models to extract informative features from visual, audio, and textual modalities. Specifically, for the visual modality, we design a dual-branch visual encoder that captures both global frame-level features and localized facial representations. For the textual modality, we introduce a context-enriched method that employs large language models to enrich emotional cues within the input text. To effectively integrate these multimodal features, we propose a fusion strategy comprising two key components, i.e., self-attention mechanisms for dynamic modality weighting, and residual connections to preserve original representations. Beyond architectural design, we further refine noisy labels in the training set by a multi-source labeling strategy. Our approach achieves a substantial performance improvement over the official baseline on the MER2025-SEMI dataset, attaining a weighted F-score of 87.49% compared to 78.63%, thereby validating the effectiveness of the proposed framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。