arXiv:2511.01444cs.AI2025-11被引 8

通过双信息瓶颈提升多模态情感分析的抗噪能力与融合效果

Robust Multimodal Sentiment Analysis via Double Information Bottleneck

论文配图:Robust Multimodal Sentiment Analysis via Double Information Bottleneck
图 1 · 摘自论文原文
  • 采用双信息瓶颈机制,分别压缩单模态和融合多模态表示
  • 在CMU-MOSI上达47.4%准确率,噪声下性能下降仅0.36%
  • 适合处理含噪数据或需鲁棒融合的多模态任务

多模态情感分析在多个研究领域备受关注。尽管算法设计取得进展,现有方法仍存在两大缺陷:单模态数据受噪声污染导致跨模态交互失真,以及多模态表示融合不充分,造成判别性信息丢失而冗余信息保留。为此,本文提出双信息瓶颈(DIB)策略,构建强健的统一紧凑多模态表示。DIB基于低秩Renyi熵函数实现,相较传统Shannon熵方法具备更强抗噪能力与高维数据计算可行性。其包含两个核心模块:1)通过最大化任务相关性并剔除冗余信息,压缩单模态数据表示;2)引入新型注意力瓶颈融合机制,保障多模态表示的判别能力。实验在CMU-MOSI、CMU-MOSEI、CH-SIMS和MVSA-Single上验证,模型在CMU-MOSI上达47.4%准确率(Acc-7),CH-SIMS上获81.63% F1-score,优于次优基线1.19%;在噪声环境下,于CMU-MOSI与CMU-MOSEI上性能仅下降0.36%和0.29%。

原文摘要 · Abstract (English)

Multimodal sentiment analysis has received significant attention across diverse research domains. Despite advancements in algorithm design, existing approaches suffer from two critical limitations: insufficient learning of noise-contaminated unimodal data, leading to corrupted cross-modal interactions, and inadequate fusion of multimodal representations, resulting in discarding discriminative unimodal information while retaining multimodal redundant information. To address these challenges, this paper proposes a Double Information Bottleneck (DIB) strategy to obtain a powerful, unified compact multimodal representation. Implemented within the framework of low-rank Renyi's entropy functional, DIB offers enhanced robustness against diverse noise sources and computational tractability for high-dimensional data, as compared to the conventional Shannon entropy-based methods. The DIB comprises two key modules: 1) learning a sufficient and compressed representation of individual unimodal data by maximizing the task-relevant information and discarding the superfluous information, and 2) ensuring the discriminative ability of multimodal representation through a novel attention bottleneck fusion mechanism. Consequently, DIB yields a multimodal representation that effectively filters out noisy information from unimodal data while capturing inter-modal complementarity. Extensive experiments on CMU-MOSI, CMU-MOSEI, CH-SIMS, and MVSA-Single validate the effectiveness of our method. The model achieves 47.4% accuracy under the Acc-7 metric on CMU-MOSI and 81.63% F1-score on CH-SIMS, outperforming the second-best baseline by 1.19%. Under noise, it shows only 0.36% and 0.29% performance degradation on CMU-MOSI and CMU-MOSEI respectively.

多模态情感分析信息瓶颈抗噪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。