解决多模态情感识别中缺失数据问题,提升模型鲁棒性。
C$^2$MOE: Consistency and Complementarity-guided Mixture of Experts for Incomplete Multimodal Emotion Learning

- 分离模态一致性与互补性,通过专家网络实现联合学习
- 在多种缺失场景下性能超越现有方法,准确率显著提升
- 适合处理真实场景中不完整多模态数据的研究者
近年来,对话中的多模态情感识别(MERC)依赖完整的多模态输入。然而,现实数据常因传输错误或用户行为导致模态缺失,严重降低模型性能。现有方法虽通过跨模态一致性学习提升鲁棒性,但忽视模态互补性,造成重建偏差。为此,本文提出C²MOE:一种基于一致性和互补性引导的混合专家框架,用于不完整多模态情感学习。该方法在信息论框架下统一表示学习与缺失模态补全。具体地,通过感知交互的专家将多模态知识分解为一致性与互补性成分:一致性通过最大化跨模态可预测性捕捉,互补性通过最大化模态间条件熵保留。在此基础上,C²MOE设计双分支预测机制,在模态缺失时实现鲁棒补全:一致性分支通过最小化不确定性对齐补全特征与联合分布,互补性分支通过熵最大化利用模态特有线索。最后,引入可学习重加权模块,动态分配各专家输出权重,实现鲁棒自适应融合。在多个MERC基准上的大量实验表明,C²MOE在不同缺失设置下均持续优于现有先进方法,验证了其鲁棒性与泛化能力。
原文摘要 · Abstract (English)
Recent advances in Multimodal Emotion Recognition in Conversations (MERC) highlight its reliance on complete multimodal inputs. However, real-world data often suffer from missing modalities due to transmission errors or user behavior, severely degrading model performance. Existing methods enhance robustness via cross-modal consistency learning but largely ignore modality complementarity, leading to biased reconstructions. To address this limitation, we propose C$2$MOE, a novel Consistency and Complementarity-guided Mixture of Experts framework for incomplete multimodal emotion learning. Our approach unifies representation learning and missing modality imputation within a principled information-theoretic framework. Specifically, multimodal knowledge is factorized into consistency and complementarity components via interaction-aware experts. Consistency is captured by maximizing cross-modal predictability, while complementarity is preserved by maximizing conditional entropy between modalities. Building upon this decomposition, C$2$MOE introduces a dual-branch prediction mechanism for robust imputation under missing modalities. The consistency branch aligns imputed features with the joint distribution by minimizing uncertainty, and the complementarity branch exploits modality-unique cues via entropy maximization. Finally, C$2$MOE employs a learnable reweighting module that dynamically assigns importance scores to each expert's output, yielding a robust and adaptive fusion for imputation. Extensive experiments on multiple MERC benchmarks demonstrate that C$2$MOE consistently surpasses state-of-the-art methods across various missing-modality settings, validating its robustness and generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。