针对多模态情感分析中数据缺失问题,提出以语言为主导的鲁棒学习方法。
Towards Robust Multimodal Sentiment Analysis with Incomplete Data
- 以语言模态为主导,设计抗噪学习网络提升鲁棒性
- 在MOSI、MOSEI等数据集上显著优于现有方法
- 适用于语言信息丰富但其他模态缺失的场景
多模态情感分析(MSA)领域正关注数据不完整问题。鉴于语言模态通常包含密集的情感信息,本文将其视为主导模态,提出一种语言主导的抗噪学习网络(LNLN),实现鲁棒的MSA。LNLN包含主导模态修正(DMC)模块和基于主导模态的多模态学习(DMML)模块,通过保障主导模态表示质量,增强模型在多种噪声情景下的鲁棒性。此外,在多个主流数据集(如MOSI、MOSEI、SIMS)上,采用多样化且有意义的随机缺失设置进行全面实验,相较现有评估更具统一性、透明性和公平性。实验证明,LNLN在诸多挑战性评价指标下持续超越基线方法,表现优异。
原文摘要 · Abstract (English)
The field of Multimodal Sentiment Analysis (MSA) has recently witnessed an emerging direction seeking to tackle the issue of data incompleteness. Recognizing that the language modality typically contains dense sentiment information, we consider it as the dominant modality and present an innovative Language-dominated Noise-resistant Learning Network (LNLN) to achieve robust MSA. The proposed LNLN features a dominant modality correction (DMC) module and dominant modality based multimodal learning (DMML) module, which enhances the model's robustness across various noise scenarios by ensuring the quality of dominant modality representations. Aside from the methodical design, we perform comprehensive experiments under random data missing scenarios, utilizing diverse and meaningful settings on several popular datasets (\textit{e.g.,} MOSI, MOSEI, and SIMS), providing additional uniformity, transparency, and fairness compared to existing evaluations in the literature. Empirically, LNLN consistently outperforms existing baselines, demonstrating superior performance across these challenging and extensive evaluation metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。