arXiv:2603.25752cs.CLcs.SD2026-03

针对多模态情感识别中噪声与模态失衡问题,提出关系感知的去噪融合模型。

Relational graph-driven differential denoising and diffusion attention fusion for multimodal conversation emotion recognition

  • 通过差分Transformer显式计算注意力图差异,抑制噪声并增强时序一致性。
  • 构建模态内与跨模态关系子图,细粒度建模说话人情绪依赖关系。
  • 引入文本引导的跨模态扩散机制,提升融合的语义一致性与鲁棒性。

在真实场景中,音频和视频信号常受环境噪声及采集条件限制,导致提取特征包含大量噪声。同时,不同模态间存在数据质量与信息承载能力的不平衡,二者共同造成融合阶段的信息失真与权重偏差,影响整体识别性能。现有方法多忽略噪声模态的影响,依赖隐式加权建模模态重要性,未能显式体现文本模态在情感理解中的主导作用。为此,本文提出一种关系感知的去噪与扩散注意力融合模型(Relation-aware Denoising and Diffusion Attention Fusion, R-DDAF)用于多模态对话情感识别(MCER)。首先,设计差分Transformer,显式计算两个注意力图的差异,增强时间一致信息,抑制非时序噪声,实现对音频与视频模态的有效去噪。其次,构建模态特定与跨模态关系子图,捕捉说话人依赖的情绪关联,实现对模态内与跨模态关系的细粒度建模。最后,引入文本引导的跨模态扩散机制,利用自注意力建模模态内依赖,并自适应地将音视频信息扩散至文本流中,确保更鲁棒、语义对齐的多模态融合。

原文摘要 · Abstract (English)

In real-world scenarios, audio and video signals are often subject to environmental noise and limited acquisition conditions, resulting in extracted features containing excessive noise. Furthermore, there is an imbalance in data quality and information carrying capacity between different modalities. These two issues together lead to information distortion and weight bias during the fusion phase, impairing overall recognition performance. Most existing methods neglect the impact of noisy modalities and rely on implicit weighting to model modality importance, thereby failing to explicitly account for the predominant contribution of the textual modality in emotion understanding. To address these issues, we propose a relation-aware denoising and diffusion attention fusion model for MCER. Specifically, we first design a differential Transformer that explicitly computes the differences between two attention maps, thereby enhancing temporally consistent information while suppressing time-irrelevant noise, which leads to effective denoising in both audio and video modalities. Second, we construct modality-specific and cross-modality relation subgraphs to capture speaker-dependent emotional dependencies, enabling fine-grained modeling of intra- and inter-modal relationships. Finally, we introduce a text-guided cross-modal diffusion mechanism that leverages self-attention to model intra-modal dependencies and adaptively diffuses audiovisual information into the textual stream, ensuring more robust and semantically aligned multimodal fusion.

情感识别多模态融合去噪注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。