arXiv:2604.14204cs.SDcs.AI2026-04

分离跨模态冗余信息,提升对话情绪识别准确率。

Disentangled Dual-Branch Graph Learning for Conversational Emotion Recognition

  • 双空间解耦+双分支图学习,分离共性与特异性特征。
  • 在IEMOCAP和MELD上超越主流基线模型。
  • 适合研究多模态情感分析与高阶交互建模的学者。

对话中的多模态情绪识别旨在通过联合建模文本、语音和视觉线索来推断话语级情绪。尽管近期取得进展,仍面临跨模态冗余信息、语义对齐不充分以及高阶说话人交互建模不足等挑战。为此,我们提出一种结合双空间特征解耦与双分支图学习的框架。采用共享编码器和模态特定编码器分离出模态不变与模态特异性表示。不变特征由傅里叶图神经网络建模,捕捉全局一致性与互补模式,并引入频域对比目标以增强判别力。同时,在模态特异性特征上构建说话人感知超图,以建模高阶交互关系,并施加说话人一致性约束以保持语义连贯性。最后融合两分支进行话语级情绪预测。在IEMOCAP和MELD数据集上的实验表明,该方法显著优于强基线,验证了其有效性。

原文摘要 · Abstract (English)

Multimodal emotion recognition in conversations aims to infer utterance-level emotions by jointly modeling textual, acoustic, and visual cues within context. Despite recent progress, key challenges remain, including redundant cross-modal information, imperfect semantic alignment, and insufficient modeling of high-order speaker interactions. To address these issues, we propose a framework that combines dual-space feature disentanglement with dual-branch graph learning. A shared encoder and modality-specific encoders are used to separate modality-invariant and modality-specific representations. The invariant features are modeled by a Fourier graph neural network to capture global consistency and complementary patterns, with a frequency-domain contrastive objective to enhance discriminability. In parallel, a speaker-aware hypergraph is constructed over modality-specific features to model high-order interactions, along with a speaker-consistency constraint to maintain coherent semantics. Finally, the two branches are fused for utterance-level emotion prediction. Experiments on IEMOCAP and MELD demonstrate that the proposed method achieves superior performance over strong baselines, validating its effectiveness.

情绪识别多模态图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。