让多模态情感识别更准:动态加权不同模态,自动处理信号不确定性。
EmoEUS: Uncertainty Supervision for Multimodal Emotion Recognition in Conversation

- 用学习到的方差估计动态调整各模态权重,实现不确定性感知融合。
- 在IEMOCAP和MELD数据集上准确率超越现有最优方法。
- 适合需要高鲁棒性情感分析的应用场景,如对话系统、心理评估。
对话中的多模态情感识别(MERC)可通过融合多源信息提升性能,但现有融合方法常忽略因矛盾线索、噪声波动或模态缺失导致的模态特异性不确定性。本文提出EmoEUS框架,通过显式不确定性监督实现更优的多模态融合。该方法利用学习到的方差估计动态加权各模态,并引入显式损失函数,使每个话语的预测方差与其分布表示与情绪-模态特定聚类中心之间的距离对齐。在IEMOCAP和MELD数据集上的实验表明,EmoEUS持续优于当前最佳方法。
原文摘要 · Abstract (English)
Multimodal emotion recognition in conversation (MERC) can leverage multimodal and contextual cues to boost recognition performance. However, existing fusion approaches in MERC often ignore modality-specific uncertainty across utterances caused by conflicting cues, varying noise, and missing modality-specific signals. We propose EmoEUS, an explicit uncertainty supervision framework for MERC. EmoEUS performs uncertainty-aware multimodal fusion by dynamically weighting modalities using learned variance estimates. We also introduce an explicitly supervised loss that aligns each utterance's predicted variance with the distance between the utterance's distributional representation and its emotion- and modality-specific cluster center. Experiments on IEMOCAP and MELD show that EmoEUS consistently outperforms state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。