arXiv:2412.03625cs.CL2024-12被引 8

用BERT+ResNet融合图文特征,提升情感分析准确率。

Multimodal Sentiment Analysis Based on BERT and ResNet

  • 结合BERT与ResNet分别提取文本和图像特征
  • 采用注意力机制融合多模态信息,准确率达74.5%
  • 适合需要跨模态分析的NLP与CV研究者

随着互联网和社交媒体的快速发展,文本与图像等多模态数据在情感分析中日益重要。然而,现有方法难以有效融合文本与图像特征,限制了分析精度。为此,提出一种结合BERT与ResNet的多模态情感分析框架。BERT在自然语言处理中表现出强大的文本表征能力,ResNet在计算机视觉领域具备优秀的图像特征提取性能。首先,利用BERT提取文本特征向量,使用ResNet提取图像特征表示;随后,探索多种特征融合策略,最终选择基于注意力机制的融合模型,充分挖掘文本与图像间的互补信息。在公开数据集MAVA-single上的实验结果表明,相较于仅使用BERT或ResNet的单模态模型,所提多模态模型在准确率和F1分数上均有提升,最高准确率达到74.5%。该研究为多模态情感分析提供了新思路,并展示了BERT与ResNet在跨域融合中的应用潜力。未来将探索更先进的特征融合技术与优化策略,以进一步提升模型精度与泛化能力。

原文摘要 · Abstract (English)

With the rapid development of the Internet and social media, multi-modal data (text and image) is increasingly important in sentiment analysis tasks. However, the existing methods are difficult to effectively fuse text and image features, which limits the accuracy of analysis. To solve this problem, a multimodal sentiment analysis framework combining BERT and ResNet was proposed. BERT has shown strong text representation ability in natural language processing, and ResNet has excellent image feature extraction performance in the field of computer vision. Firstly, BERT is used to extract the text feature vector, and ResNet is used to extract the image feature representation. Then, a variety of feature fusion strategies are explored, and finally the fusion model based on attention mechanism is selected to make full use of the complementary information between text and image. Experimental results on the public dataset MAVA-single show that compared with the single-modal models that only use BERT or ResNet, the proposed multi-modal model improves the accuracy and F1 score, reaching the best accuracy of 74.5%. This study not only provides new ideas and methods for multimodal sentiment analysis, but also demonstrates the application potential of BERT and ResNet in cross-domain fusion. In the future, more advanced feature fusion techniques and optimization strategies will be explored to further improve the accuracy and generalization ability of multimodal sentiment analysis.

情感分析多模态BERTResNet

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。