arXiv:2503.07943cs.CVcs.CL2025-03被引 2

用BERT+DINOv2融合图文信息,提升情感分析准确性

Enhancing Sentiment Analysis through Multimodal Fusion: A BERT-DINOv2 Approach

  • 结合BERT与DINOv2提取文本与图像特征
  • 三种注意力融合机制在三个数据集上均超越基线
  • 适合需要多模态情感理解的研究与应用

多模态情感分析通过整合文本、图像、音频等多源信息,突破传统仅依赖文本的局限。本文提出一种融合文本与图像的新型多模态情感分析架构:采用BERT进行文本特征提取,DINOv2作为视觉编码器处理图像特征,并设计三种融合策略——基础融合、自注意力融合与双注意力融合。在Memotion 7k、MVSA single和MVSA multi三个数据集上的实验验证了该架构的有效性,显著提升了情感分类性能。

原文摘要 · Abstract (English)

Multimodal sentiment analysis enhances conventional sentiment analysis, which traditionally relies solely on text, by incorporating information from different modalities such as images, text, and audio. This paper proposes a novel multimodal sentiment analysis architecture that integrates text and image data to provide a more comprehensive understanding of sentiments. For text feature extraction, we utilize BERT, a natural language processing model. For image feature extraction, we employ DINOv2, a vision-transformer-based model. The textual and visual latent features are integrated using proposed fusion techniques, namely the Basic Fusion Model, Self Attention Fusion Model, and Dual Attention Fusion Model. Experiments on three datasets, Memotion 7k dataset, MVSA single dataset, and MVSA multi dataset, demonstrate the viability and practicality of the proposed multimodal architecture.

情感分析多模态BERTDINOv2

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。