arXiv:2412.09784cs.CLcs.AI2024-12被引 13

通过自训练动态融合模态内与模态间交互,提升少标注场景下的情感分析性能。

Semi-IIN: Semi-supervised Intra-inter modal Interaction Learning Network for Multimodal Sentiment Analysis

  • 分步捕捉模态内/外交互,用门控机制动态选择有效信息
  • 在MOSI和MOSEI数据集上达到新SOTA,显著优于现有方法
  • 适用于标注稀缺但有大量无标签数据的多模态情感分析场景

尽管多模态情感分析是一个富有前景的研究方向,但现有方法依赖高成本标注,且面临标签模糊问题,难以获取高质量标注数据。此外,不同样本中模态内或模态间交互的重要性各异,选择合适的交互方式至关重要。为此,我们提出Semi-IIN:一种半监督的模态内-模态间交互学习网络。该模型结合掩码注意力与门控机制,在独立捕获模态内与模态间交互信息后,实现有效的动态选择。结合自训练策略,充分挖掘无标签数据中的知识。在两个公开数据集MOSI和MOSEI上的实验表明,Semi-IIN在多个指标上表现优异,刷新了当前最佳性能。代码已开源:https://github.com/flow-ljh/Semi-IIN。

原文摘要 · Abstract (English)

Despite multimodal sentiment analysis being a fertile research ground that merits further investigation, current approaches take up high annotation cost and suffer from label ambiguity, non-amicable to high-quality labeled data acquisition. Furthermore, choosing the right interactions is essential because the significance of intra- or inter-modal interactions can differ among various samples. To this end, we propose Semi-IIN, a Semi-supervised Intra-inter modal Interaction learning Network for multimodal sentiment analysis. Semi-IIN integrates masked attention and gating mechanisms, enabling effective dynamic selection after independently capturing intra- and inter-modal interactive information. Combined with the self-training approach, Semi-IIN fully utilizes the knowledge learned from unlabeled data. Experimental results on two public datasets, MOSI and MOSEI, demonstrate the effectiveness of Semi-IIN, establishing a new state-of-the-art on several metrics. Code is available at https://github.com/flow-ljh/Semi-IIN.

多模态情感分析半监督自训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。