arXiv:2503.23721cs.LGcs.AI2025-03被引 5

用单模态教师指导多模态情感识别,提升小众情绪识别效果

Unimodal-driven Distillation in Multimodal Emotion Recognition with Dynamic Fusion

  • 通过动态专家混合捕捉跨模态细粒度交互
  • 在IEMOCAP和MELD上超越现有方法,尤其改善少数类情绪识别
  • 适合需要精准识别复杂情绪的对话系统研究者

对话中的多模态情感识别(MERC)旨在融合文本、音频和视频信息以识别情绪,对智能对话系统与观点分析至关重要。现有方法直接进行异构模态融合,常因模态差异大且缺乏指导而出现学习偏差。本文提出SUMMER框架,结合分层跨模态融合与交互式知识蒸馏,利用预训练单模态教师在潜在空间与输出空间引导融合。核心组件包括稀疏动态专家混合(SDMoE)实现细粒度交互,分层跨模态融合(HCMF)有效整合异构模态,以及交互式知识蒸馏(IKD)。在IEMOCAP和MELD数据集上的实验表明,SUMMER显著优于当前最优方法,尤其在识别少数类及语义相似情绪时表现突出。

原文摘要 · Abstract (English)

Multimodal Emotion Recognition in Conversations (MERC) identifies emotional states across text, audio and video, which is essential for intelligent dialogue systems and opinion analysis. Existing methods emphasize heterogeneous modal fusion directly for cross-modal integration, but often suffer from disorientation in multimodal learning due to modal heterogeneity and lack of instructive guidance. In this work, we propose SUMMER, a novel heterogeneous multimodal integration framework leveraging Mixture of Experts with Hierarchical Cross-modal Fusion and Interactive Knowledge Distillation. Key components include a Sparse Dynamic Mixture of Experts (SDMoE) for capturing dynamic token-wise interactions, a Hierarchical Cross-Modal Fusion (HCMF) for effective fusion of heterogeneous modalities, and Interactive Knowledge Distillation (IKD), which uses a pre-trained unimodal teacher to guide multimodal fusion in latent and logit spaces. Experiments on IEMOCAP and MELD show SUMMER outperforms state-of-the-art methods, particularly in recognizing minority and semantically similar emotions.

情感识别多模态融合知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。