arXiv:2409.14703cs.LGcs.CL2024-09EMNLP被引 54

用CLIP提升多模态迷因分类,同时识别仇恨、立场和幽默。

MemeCLIP: Leveraging CLIP Representations for Multimodal Meme Classification

  • 基于预训练CLIP构建新框架MemeCLIP,保留原始模型知识。
  • 在5063张酷儿骄傲迷因上实现比现有方法更优的多任务表现。
  • 适合关注社会情绪分析与多模态理解的研究者使用。

文本嵌入图像的复杂性对机器学习构成挑战,需综合理解其多维度表达。现有研究多聚焦单一方面(如仇恨言论),本文扩展至仇恨、仇恨对象、立场和幽默四个语言层面。提出新数据集PrideMM,包含5,063张与酷儿骄傲运动相关的文本嵌入图像,填补资源空白。在PrideMM上测试单模态与多模态基线方法,建立各任务基准。提出MemeCLIP框架,实现高效下游学习并保留预训练CLIP知识。实验表明,MemeCLIP在两个真实数据集上优于已有框架。进一步对比MemeCLIP与零样本GPT-4在仇恨分类上的表现。通过误分类样例定性分析揭示模型局限。代码与数据集已开源。

原文摘要 · Abstract (English)

The complexity of text-embedded images presents a formidable challenge in machine learning given the need for multimodal understanding of multiple aspects of expression conveyed by them. While previous research in multimodal analysis has primarily focused on singular aspects such as hate speech and its subclasses, this study expands this focus to encompass multiple aspects of linguistics: hate, targets of hate, stance, and humor. We introduce a novel dataset PrideMM comprising 5,063 text-embedded images associated with the LGBTQ+ Pride movement, thereby addressing a serious gap in existing resources. We conduct extensive experimentation on PrideMM by using unimodal and multimodal baseline methods to establish benchmarks for each task. Additionally, we propose a novel framework MemeCLIP for efficient downstream learning while preserving the knowledge of the pre-trained CLIP model. The results of our experiments show that MemeCLIP achieves superior performance compared to previously proposed frameworks on two real-world datasets. We further compare the performance of MemeCLIP and zero-shot GPT-4 on the hate classification task. Finally, we discuss the shortcomings of our model by qualitatively analyzing misclassified samples. Our code and dataset are publicly available at: https://github.com/SiddhantBikram/MemeCLIP.

多模态迷因分类CLIP社会情感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。