用图像物体名称增强文本与图像情感分析的融合效果
Leveraging Textual-Cues for Enhancing Multimodal Sentiment Analysis by Object Recognition
- 通过识别图像中的物体并提取其名称,与文本结合提升情感分析
- 在两个数据集上,联合使用物体名称使整体情感判断准确率提升
- 适合关注多模态情感分析与视觉语义融合的研究者
多模态情感分析需同时处理图像和文本数据,面临模态差异、情感模糊和上下文复杂等挑战。本文在两个数据集上实验了单独及联合分析图像与文本情感的方法。提出一种基于物体识别的新方法TEMSA(Textual-Cues for Enhancing Multimodal Sentiment Analysis),将图像中检测到的所有物体名称与对应文本结合,形成融合数据(TEMS)。实验表明,仅使用所有物体名称进行联合分析时,整体多模态情感判断性能优于单独分析。该研究推动了多模态情感分析的发展,验证了TEMSA在融合图像与文本信息方面的有效性。
原文摘要 · Abstract (English)
Multimodal sentiment analysis, which includes both image and text data, presents several challenges due to the dissimilarities in the modalities of text and image, the ambiguity of sentiment, and the complexities of contextual meaning. In this work, we experiment with finding the sentiments of image and text data, individually and in combination, on two datasets. Part of the approach introduces the novel `Textual-Cues for Enhancing Multimodal Sentiment Analysis' (TEMSA) based on object recognition methods to address the difficulties in multimodal sentiment analysis. Specifically, we extract the names of all objects detected in an image and combine them with associated text; we call this combination of text and image data TEMS. Our results demonstrate that only TEMS improves the results when considering all the object names for the overall sentiment of multimodal data compared to individual analysis. This research contributes to advancing multimodal sentiment analysis and offers insights into the efficacy of TEMSA in combining image and text data for multimodal sentiment analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。