自适应门控融合提升多模态情感分析鲁棒性
Beyond Simple Fusion: Adaptive Gated Fusion for Robust Multimodal Sentiment Analysis
- 用双门控机制按信息熵和模态重要性动态加权特征
- 在CMU-MOSI/MOSEI上准确率显著优于基线模型
- 适合处理噪声、缺失或冲突的多模态数据
多模态情感分析通过融合文本、音频、视觉等模态信息提升情感预测效果。然而,简单融合方法无法应对模态质量差异,如噪声、缺失或语义冲突,导致在识别细微情绪时性能下降。为此,我们提出一种高效且简单的自适应门控融合网络(AGFN),基于信息熵和模态重要性设计双门控融合机制,在单模态编码与跨模态交互后动态调整特征权重,有效抑制噪声模态影响,优先利用关键线索。在CMU-MOSI和CMU-MOSEI数据集上的实验表明,AGFN显著优于强基线,在准确率上表现优异,能有效分辨细微情绪。特征表示可视化分析显示,AGFN通过降低特征位置与预测误差的相关性,扩大特征分布范围,减少对特定位置的依赖,从而构建更具泛化能力的多模态特征表示。
原文摘要 · Abstract (English)
Multimodal sentiment analysis (MSA) leverages information fusion from diverse modalities (e.g., text, audio, visual) to enhance sentiment prediction. However, simple fusion techniques often fail to account for variations in modality quality, such as those that are noisy, missing, or semantically conflicting. This oversight leads to suboptimal performance, especially in discerning subtle emotional nuances. To mitigate this limitation, we introduce a simple yet efficient \textbf{A}daptive \textbf{G}ated \textbf{F}usion \textbf{N}etwork that adaptively adjusts feature weights via a dual gate fusion mechanism based on information entropy and modality importance. This mechanism mitigates the influence of noisy modalities and prioritizes informative cues following unimodal encoding and cross-modal interaction. Experiments on CMU-MOSI and CMU-MOSEI show that AGFN significantly outperforms strong baselines in accuracy, effectively discerning subtle emotions with robust performance. Visualization analysis of feature representations demonstrates that AGFN enhances generalization by learning from a broader feature distribution, achieved by reducing the correlation between feature location and prediction error, thereby decreasing reliance on specific locations and creating more robust multimodal feature representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。