用大模型自动标注海量梗图,提升文化语义理解能力
Large Vision-Language Models for Knowledge-Grounded Data Annotation of Memes
- 基于大视觉语言模型构建自动化标注流程
- 建成包含3.3万张梗图的50模板数据集
- 适配梗图检索的CLIP模型显著提升匹配精度
梗图作为融合视觉与文本的传播形式,承载幽默、讽刺与文化信息。现有研究多集中于情绪分类、生成与传播分析,却忽视深层语义理解与图文匹配。为此,本文构建了经典梗图-50模板(CM50)数据集,包含超过3.3万张以50个流行模板为核心的梗图。提出一种基于大视觉语言模型的自动化知识引导标注流程,可高效生成高质量图像描述、梗图标题及修辞手法标签,显著降低人工标注成本。此外,设计了专用于梗图-文本检索的mtrCLIP模型,通过跨模态嵌入增强匹配性能。贡献包括:(1)首个大规模梗图研究数据集;(2)可扩展的自动化标注框架;(3)经微调的面向梗图检索的CLIP模型,全面推动梗图的规模化理解与分析。
原文摘要 · Abstract (English)
Memes have emerged as a powerful form of communication, integrating visual and textual elements to convey humor, satire, and cultural messages. Existing research has focused primarily on aspects such as emotion classification, meme generation, propagation, interpretation, figurative language, and sociolinguistics, but has often overlooked deeper meme comprehension and meme-text retrieval. To address these gaps, this study introduces ClassicMemes-50-templates (CM50), a large-scale dataset consisting of over 33,000 memes, centered around 50 popular meme templates. We also present an automated knowledge-grounded annotation pipeline leveraging large vision-language models to produce high-quality image captions, meme captions, and literary device labels overcoming the labor intensive demands of manual annotation. Additionally, we propose a meme-text retrieval CLIP model (mtrCLIP) that utilizes cross-modal embedding to enhance meme analysis, significantly improving retrieval performance. Our contributions include:(1) a novel dataset for large-scale meme study, (2) a scalable meme annotation framework, and (3) a fine-tuned CLIP for meme-text retrieval, all aimed at advancing the understanding and analysis of memes at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。