arXiv:2608.20019cs.AI2026-08

解决跨模态情感分析中未见组合的泛化难题,提升模型鲁棒性。

Contrastive Mixed Prompt Learning for Incomplete Multimodal Sentiment Analysis with Unseen Modality Combination

论文配图:Contrastive Mixed Prompt Learning for Incomplete Multimodal Sentiment Analysis with Unseen Modality Combination
图 1 · 摘自论文原文
  • 设计对比混合提示学习框架,通过对比特征与软路由机制增强跨模态表征。
  • 在三个数据集上相比最优方法准确率提升超5%。
  • 特别适合处理训练时未见过的模态组合场景,如视频+文本缺失的现实应用。

不完整多模态情感分析近年来受到广泛关注。现有方法通常假设数据随机缺失或仅针对特定缺失模式设计,忽略了训练与测试阶段模态组合不一致的问题。然而在真实场景中,测试阶段常遇到训练中未出现的模态组合,导致模型泛化能力不足、性能不稳定。本文提出不完整多模态情感分析中未见模态组合问题(IMSAUMC),旨在提升模型对未见模态组合的泛化能力。为此,我们提出对比混合提示学习模型(CMPL),引入标签引导的对比特征学习机制以获取鲁棒且具有判别性的跨模态表示;设计带有软路由的模态组合提示,促进对多种模态组合的学习;并提出三种提示对比学习策略,有效学习未见模态组合对应的提示,显著增强模型在多样化测试场景下的泛化能力。在三个常用数据集上的大量实验表明,CMPL相比当前最优方法准确率提升超过5%。

原文摘要 · Abstract (English)

Incomplete multimodal sentiment analysis has garnered significant attention in recent years. Existing approaches typically assume that data is missing at random or are designed specifically for certain missing patterns, ignoring the modality combination inconsistency between training and testing phases. However, in real-world scenarios, the testing phase often encounters modal combinations that were not present during the training phase, which leads to insufficient generalization capabilities and unstable performance. In this paper, we introduce the problem of Incomplete Multimodal Sentiment Analysis with Unseen Modality Combinations (IMSAUMC), aiming to enhance model generalization for unseen modality combinations. To address this challenge, we propose the model named $\textbf{C}$ontrastive $\textbf{M}$ixed $\textbf{P}$rompt $\textbf{L}$earning ($\textsf{CMPL}$) for IMSAUMC. It introduces a label-guided contrastive feature learning mechanism to learn robust and discriminative cross-modal representations. Additionally, we design modality-combination prompts with a soft router to facilitate better learning of various modality combinations. Furthermore, we introduce three prompt contrastive learning strategies, which enable effective learning of prompts corresponding to unseen modality combinations, thereby significantly strengthening the model's generalization capabilities in diverse testing scenarios. Extensive experiments on three widely used datasets demonstrate that $\textsf{CMPL}$ achieves more than a 5% improvement in accuracy compared to state-of-the-art approaches.

多模态情感分析提示学习泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。