arXiv:2603.10043cs.MMcs.AI2026-03被引 1

提出动态图网络解决多模态情感识别中噪声干扰与模态失衡问题。

AMB-DSGDN: Adaptive Modality-Balanced Dynamic Semantic Graph Differential Network for Multimodal Emotion Recognition

  • 构建跨说话人与自连续的子图,捕捉情感依赖关系。
  • 用差异注意力机制消除共性噪声,提升表示纯净度。
  • 自适应调制模态丢弃率,平衡语音、视觉等非主导模态贡献。

多模态对话情感识别通过融合文本、视觉和音频模态捕捉情绪线索。然而现有方法在建模情感依赖关系和学习多模态表征方面仍存在明显局限:一方面难以有效过滤多模态特征中的冗余或噪声信号,阻碍了说话人内外情感状态动态演变的准确捕捉;另一方面,在多模态特征学习中,主导模态常压制非主导模态(如语音与视觉)的互补贡献,制约整体识别性能。为此,本文提出自适应模态平衡动态语义图微分网络(AMB-DSGDN)。首先为文本、语音、视觉分别构建模态专属子图,每类子图包含说话人内与跨说话人图以捕捉自连续性和跨说话人情感依赖。在此基础上引入差异图注意力机制,计算两组注意力分布间的差异,通过显式对比,消除共性噪声模式,保留模态特异性和上下文相关信号,从而获得更纯净、更具判别性的情感表示。此外,设计自适应模态平衡机制,根据各模态在情感建模中的相对贡献动态估计其丢弃概率。

原文摘要 · Abstract (English)

Multimodal dialogue emotion recognition captures emotional cues by fusing text, visual, and audio modalities. However, existing approaches still suffer from notable limitations in modeling emotional dependencies and learning multimodal representations. On the one hand, they are unable to effectively filter out redundant or noisy signals within multimodal features, which hinders the accurate capture of the dynamic evolution of emotional states across and within speakers. On the other hand, during multimodal feature learning, dominant modalities tend to overwhelm the fusion process, thereby suppressing the complementary contributions of non-dominant modalities such as speech and vision, ultimately constraining the overall recognition performance. To address these challenges, we propose an Adaptive Modality-Balanced Dynamic Semantic Graph Differential Network (AMB-DSGDN). Concretely, we first construct modality-specific subgraphs for text, speech, and vision, where each modality contains intra-speaker and inter-speaker graphs to capture both self-continuity and cross-speaker emotional dependencies. On top of these subgraphs, we introduce a differential graph attention mechanism, which computes the discrepancy between two sets of attention maps. By explicitly contrasting these attention distributions, the mechanism cancels out shared noise patterns while retaining modality-specific and context-relevant signals, thereby yielding purer and more discriminative emotional representations. In addition, we design an adaptive modality balancing mechanism, which estimates a dropout probability for each modality according to its relative contribution in emotion modeling.

多模态情感识别图神经网络注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。