用上下文注意力融合文本与图像,提升灾害舆情分析准确率。
Contextual Attention-Based Multimodal Fusion of LLM and CNN for Sentiment Analysis
- 结合CNN图像处理与LLM文本理解,引入上下文注意力机制融合多模态信息。
- 在CrisisMMD数据集上,准确率提升2.43%,F1-score提升5.18%。
- 适合需要实时灾情舆情分析的应急管理系统使用。
本文提出一种新型多模态情感分析方法,针对自然灾害背景下社交媒体舆情分析需求,旨在提升危机管理效率。不同于传统分别处理文本与图像的方法,本模型将基于卷积神经网络(CNN)的图像分析与基于大语言模型(LLM)的文本处理相结合,利用生成式预训练变换器(GPT)与提示工程从CrisisMMD数据集提取情感相关特征。为有效建模跨模态关系,引入上下文注意力机制,通过注意力层捕捉文本与视觉数据间的复杂交互。深度神经网络从融合特征中学习,显著优于现有基线方法。实验结果表明,在各类自然灾害场景下,该模型在区分信息性与非信息性社交媒体内容方面表现优异,准确率提升2.43%,F1-score提升5.18%。该方法不仅在定量指标上表现突出,还提供对危机中公众情绪的深层洞察,对实时灾害响应具有重要实践价值。
原文摘要 · Abstract (English)
This paper introduces a novel approach for multimodal sentiment analysis on social media, particularly in the context of natural disasters, where understanding public sentiment is crucial for effective crisis management. Unlike conventional methods that process text and image modalities separately, our approach seamlessly integrates Convolutional Neural Network (CNN) based image analysis with Large Language Model (LLM) based text processing, leveraging Generative Pre-trained Transformer (GPT) and prompt engineering to extract sentiment relevant features from the CrisisMMD dataset. To effectively model intermodal relationships, we introduce a contextual attention mechanism within the fusion process. Leveraging contextual-attention layers, this mechanism effectively captures intermodality interactions, enhancing the model's comprehension of complex relationships between textual and visual data. The deep neural network architecture of our model learns from these fused features, leading to improved accuracy compared to existing baselines. Experimental results demonstrate significant advancements in classifying social media data into informative and noninformative categories across various natural disasters. Our model achieves a notable 2.43% increase in accuracy and 5.18% in F1-score, highlighting its efficacy in processing complex multimodal data. Beyond quantitative metrics, our approach provides deeper insight into the sentiments expressed during crises. The practical implications extend to real time disaster management, where enhanced sentiment analysis can optimize the accuracy of emergency interventions. By bridging the gap between multimodal analysis, LLM powered text understanding, and disaster response, our work presents a promising direction for Artificial Intelligence (AI) driven crisis management solutions. Keywords:
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。