解决对话中缺失模态时的情感识别问题,提升真实场景下的鲁棒性。
Federated Dialogue-Semantic Diffusion for Emotion Recognition under Incomplete Modalities
- 通过联邦学习融合各客户端的模态扩散模型,实现跨设备协同恢复缺失信息。
- 在IEMOCAP等数据集上,不同缺失模式下准确率均优于现有方法。
- 适合研究多模态情感分析与隐私保护场景下的模型部署者。
对话中的多模态情感识别(MERC)通过融合多种信号提升情感理解能力。然而,现实场景中模态缺失不可预测,严重影响现有方法性能。传统缺失模态恢复依赖完整数据训练,常在固定模态缺失等极端分布下产生语义失真。为此,我们提出联邦对话引导与语义一致扩散(FedDISC)框架,首次将联邦学习引入缺失模态恢复。通过在客户端训练模态专用扩散模型并联邦聚合,再广播至缺失对应模态的客户端,克服单客户端对模态完整的依赖。DISC-Diffusion模块利用对话图网络捕捉对话依赖,结合语义条件网络确保恢复内容与已有模态在上下文、说话人身份和语义上保持一致。我们还提出交替冻结聚合策略,周期性冻结恢复与分类模块以促进协同优化。在IEMOCAP、CMUMOSI和CMUMOSEI数据集上的实验表明,FedDISC在多种缺失模式下均实现更优的情感分类性能,显著优于现有方法。
原文摘要 · Abstract (English)
Multimodal Emotion Recognition in Conversations (MERC) enhances emotional understanding through the fusion of multimodal signals. However, unpredictable modality absence in real-world scenarios significantly degrades the performance of existing methods. Conventional missing-modality recovery approaches, which depend on training with complete multimodal data, often suffer from semantic distortion under extreme data distributions, such as fixed-modality absence. To address this, we propose the Federated Dialogue-guided and Semantic-Consistent Diffusion (FedDISC) framework, pioneering the integration of federated learning into missing-modality recovery. By federated aggregation of modality-specific diffusion models trained on clients and broadcasting them to clients missing corresponding modalities, FedDISC overcomes single-client reliance on modality completeness. Additionally, the DISC-Diffusion module ensures consistency in context, speaker identity, and semantics between recovered and available modalities, using a Dialogue Graph Network to capture conversational dependencies and a Semantic Conditioning Network to enforce semantic alignment. We further introduce a novel Alternating Frozen Aggregation strategy, which cyclically freezes recovery and classifier modules to facilitate collaborative optimization. Extensive experiments on the IEMOCAP, CMUMOSI, and CMUMOSEI datasets demonstrate that FedDISC achieves superior emotion classification performance across diverse missing modality patterns, outperforming existing approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。