解决多模态情感分析中缺失模态问题,提升模型鲁棒性
Progressive Representation Learning for Multimodal Sentiment Analysis with Incomplete Modalities
- 通过动态评估模态可靠性,识别主导模态并自适应调整融合策略
- 在三种数据集上优于现有方法,跨模态缺失场景下准确率显著提升
- 适合实际应用中存在模态缺失或噪声的复杂场景使用
多模态情感分析(MSA)旨在通过文本、语音和视觉线索综合推断人类情绪。然而,现有方法通常假设所有模态均完整,而现实应用中常因噪声、硬件故障或隐私限制导致模态缺失。不完整与完整模态间存在显著特征错位,直接融合可能破坏已学习的完整模态表示。为此,我们提出一种渐进式表示学习框架PRLF,用于不确定缺失模态条件下的MSA。PRLF引入自适应模态可靠性估计器(AMRE),利用识别置信度和Fisher信息动态量化各模态可靠性,以确定主导模态。同时,渐进交互模块(ProgInteract)迭代对齐其他模态与主导模态,增强跨模态一致性并抑制噪声。在CMU-MOSI、CMU-MOSEI和SIMS上的大量实验表明,PRLF在跨模态和同模态缺失场景下均优于当前最优方法,验证了其鲁棒性与泛化能力。
原文摘要 · Abstract (English)
Multimodal Sentiment Analysis (MSA) seeks to infer human emotions by integrating textual, acoustic, and visual cues. However, existing approaches often rely on all modalities are completeness, whereas real-world applications frequently encounter noise, hardware failures, or privacy restrictions that result in missing modalities. There exists a significant feature misalignment between incomplete and complete modalities, and directly fusing them may even distort the well-learned representations of the intact modalities. To this end, we propose PRLF, a Progressive Representation Learning Framework designed for MSA under uncertain missing-modality conditions. PRLF introduces an Adaptive Modality Reliability Estimator (AMRE), which dynamically quantifies the reliability of each modality using recognition confidence and Fisher information to determine the dominant modality. In addition, the Progressive Interaction (ProgInteract) module iteratively aligns the other modalities with the dominant one, thereby enhancing cross-modal consistency while suppressing noise. Extensive experiments on CMU-MOSI, CMU-MOSEI, and SIMS verify that PRLF outperforms state-of-the-art methods across both inter- and intra-modality missing scenarios, demonstrating its robustness and generalization capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。