arXiv:2605.21565cs.LG2026-05中稿 · Neural Computing a…

用自适应课程学习提升多模态情感识别的平衡性。

Leveraging Self-Paced Curriculum Learning for Enhanced Modality Balance in Multimodal Conversational Emotion Recognition

论文配图:Leveraging Self-Paced Curriculum Learning for Enhanced Modality Balance in Multimodal Conversational Emotion Recognition
图 1 · 摘自论文原文
  • 设计双层难度度量,分别捕捉语句和对话层面的挑战。
  • 在IEMOCAP上提升1.2%~6.6%,MELD上最高提升10.4%。
  • 可直接嵌入现有模型,适合多模态情感分析研究者。

对话中的多模态情感识别(MERC)是理解人类交互的关键任务,融合语言、面部表情和语音语调的多模态方法已取得显著进展。然而,模态错位和学习不平衡仍是主要挑战,限制了多模态信息的有效利用。为此,我们提出一种基于自适应课程学习(SPCL)的即插即用框架。引入双层难度度量,分别捕捉语句级和对话级的挑战:语句级评分建模细粒度的模态特定难度,对话级评分则捕获情绪依赖和模态一致性等更广泛的对话结构。基于这些评分,学习调度器动态引导模型从简单到复杂样本训练。将SPCL集成至现有MERC架构后,有效缓解了模态不平衡,提升了模型鲁棒性。在IEMOCAP和MELD数据集上的大量实验表明,该方法在不同架构与模态设置下均具一致性提升。在IEMOCAP上,加权F1分数相比基线提升约1.2%至6.6%;在MELD上,提升最高达10.4%。结果验证了SPCL作为轻量级通用模块在多模态情感识别中的有效性与泛化能力。

原文摘要 · Abstract (English)

Multimodal Emotion Recognition in Conversations (MERC) is a crucial task for understanding human interactions, where multimodal approaches integrating language, facial expressions, and vocal tone have achieved significant progress. However, modality misalignment and imbalanced learning remain major challenges, limiting the effective utilization of multimodal information. To address this issue, we propose a plug-and-play framework based on Self-Paced Curriculum Learning (SPCL) for MERC. We introduce a dual-level Difficulty Measurer that captures both utterance-level and conversation-level challenges. The utterance-level score models fine-grained modality-specific difficulty, while the conversation-level score captures broader dialogue structures, including emotional dependencies and modality coherence. Based on these scores, the Learning Scheduler dynamically guides training from easier to more difficult instances. By integrating SPCL into existing MERC architectures, our method alleviates modality imbalance and improves model robustness. Extensive experiments on the IEMOCAP and MELD datasets demonstrate consistent improvements across different architectures and modality settings. On IEMOCAP, SPCL improves weighted F1-score by approximately +1.2% to +6.6% over baseline models, while on MELD, gains reach up to +10.4%. These results highlight the effectiveness and generalizability of SPCL as a lightweight plug-and-play module for multimodal emotion recognition.

情感识别多模态课程学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。