arXiv:2510.23617cs.LGcs.AI2025-10中稿 · presentation at th…被引 1

用双分支注意力增强跨模态情感分析,提升文本图像融合效果

An Enhanced Dual Transformer Contrastive Network for Multimodal Sentiment Analysis

  • 先融合文本与图像特征,再通过双分支Transformer细化语义
  • 在TumEmo数据集上达78.4%准确率,优于现有方法
  • 适合做多模态情感分析、视觉语言理解的研究者参考

多模态情感分析(MSA)通过联合分析文本和图像等多源信息,实现更全面的情感理解。本文提出BERT-ViT-EF模型,采用早期融合策略结合BERT与ViT编码器,促进跨模态深度交互。进一步构建双变压器对比网络(DTCN),在文本分支增加额外的Transformer层以优化上下文表示,并引入对比学习对齐图文特征,增强多模态表征能力。在两个常用基准数据集MVSA-Single和TumEmo上的实验表明,DTCN在TumEmo上取得78.4%准确率和78.3% F1-score,在MVSA-Single上达到76.6%准确率和75.9% F1-score,验证了早期融合与深层建模的有效性。

原文摘要 · Abstract (English)

Multimodal Sentiment Analysis (MSA) seeks to understand human emotions by jointly analyzing data from multiple modalities typically text and images offering a richer and more accurate interpretation than unimodal approaches. In this paper, we first propose BERT-ViT-EF, a novel model that combines powerful Transformer-based encoders BERT for textual input and ViT for visual input through an early fusion strategy. This approach facilitates deeper cross-modal interactions and more effective joint representation learning. To further enhance the model's capability, we propose an extension called the Dual Transformer Contrastive Network (DTCN), which builds upon BERT-ViT-EF. DTCN incorporates an additional Transformer encoder layer after BERT to refine textual context (before fusion) and employs contrastive learning to align text and image representations, fostering robust multimodal feature learning. Empirical results on two widely used MSA benchmarks MVSA-Single and TumEmo demonstrate the effectiveness of our approach. DTCN achieves best accuracy (78.4%) and F1-score (78.3%) on TumEmo, and delivers competitive performance on MVSA-Single, with 76.6% accuracy and 75.9% F1-score. These improvements highlight the benefits of early fusion and deeper contextual modeling in Transformer-based multimodal sentiment analysis.

多模态情感分析Transformer对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。