用多模态方法分析网络迷因相似性与情感,准确率达67.23%
Meme Similarity and Emotion Detection using Multimodal Analysis
- 结合图像与文本嵌入,用CLIP模型计算迷因相似度
- 人机对比结果显示67.23%一致率,验证算法有效性
- 发现愤怒与喜悦是主流情绪,励志类迷因更易引发共鸣
网络迷因是在线文化的核心,融合图像与文字。现有研究多聚焦单一模态,忽视二者交互。本研究提出多模态方法,基于CLIP模型对图像和文本嵌入进行联合分析,实现跨模态迷因相似性评估。利用Reddit Meme Dataset和Memotion Dataset,提取低层视觉特征与高层语义特征,识别相似迷因对。通过50名参与者的人类判断实验,验证自动化结果,与人工判断的吻合度达67.23%。同时,采用DistilBERT模型构建文本分类器,将迷因分为六种基本情绪,结果显示愤怒与喜悦占主导,激励类迷因引发更强情绪反应。该研究推动多模态迷因分析,提升在线视觉传播与用户体验,并为平台内容治理提供支持。
原文摘要 · Abstract (English)
Internet memes are a central element of online culture, blending images and text. While substantial research has focused on either the visual or textual components of memes, little attention has been given to their interplay. This gap raises a key question: What methodology can effectively compare memes and the emotions they elicit? Our study employs a multimodal methodological approach, analyzing both the visual and textual elements of memes. Specifically, we perform a multimodal CLIP (Contrastive Language-Image Pre-training) model for grouping similar memes based on text and visual content embeddings, enabling robust similarity assessments across modalities. Using the Reddit Meme Dataset and Memotion Dataset, we extract low-level visual features and high-level semantic features to identify similar meme pairs. To validate these automated similarity assessments, we conducted a user study with 50 participants, asking them to provide yes/no responses regarding meme similarity and their emotional reactions. The comparison of experimental results with human judgments showed a 67.23\% agreement, suggesting that the computational approach aligns well with human perception. Additionally, we implemented a text-based classifier using the DistilBERT model to categorize memes into one of six basic emotions. The results indicate that anger and joy are the dominant emotions in memes, with motivational memes eliciting stronger emotional responses. This research contributes to the study of multimodal memes, enhancing both language-based and visual approaches to analyzing and improving online visual communication and user experiences. Furthermore, it provides insights for better content moderation strategies in online platforms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。