测试大模型在对话中理解梗图意图的能力,发现它们常忽略上下文。
MemeReaCon: Probing Contextual Meme Understanding in Large Vision-Language Models
- 构建含图文和评论的梗图数据集,评估模型对上下文依赖的理解。
- 主流大模型在上下文理解上表现不佳,或过度关注图像细节。
- 适合研究多模态推理、对话理解与虚假信息检测的学者使用。
梗图已成为流行的多模态网络表达形式,其含义高度依赖具体语境。现有方法多聚焦孤立分析,忽视同一梗图在不同对话中可能传递不同意图的核心挑战。为此,我们提出MemeReaCon,一个专门评估大型视觉语言模型(LVLMs)在原始语境中理解梗图的基准。从五个Reddit社区收集数据,保留图片、帖子正文及用户评论的完整上下文,并标注文本与图像协同作用、发帖者意图、梗图结构及社区反应。对主流LVLM的测试显示,模型普遍无法识别关键上下文信息,或过度关注视觉细节而忽略传播目的。MemeReaCon既可诊断当前模型缺陷,也可推动具备上下文感知能力的下一代模型发展。
原文摘要 · Abstract (English)
Memes have emerged as a popular form of multimodal online communication, where their interpretation heavily depends on the specific context in which they appear. Current approaches predominantly focus on isolated meme analysis, either for harmful content detection or standalone interpretation, overlooking a fundamental challenge: the same meme can express different intents depending on its conversational context. This oversight creates an evaluation gap: although humans intuitively recognize how context shapes meme interpretation, Large Vision Language Models (LVLMs) can hardly understand context-dependent meme intent. To address this critical limitation, we introduce MemeReaCon, a novel benchmark specifically designed to evaluate how LVLMs understand memes in their original context. We collected memes from five different Reddit communities, keeping each meme's image, the post text, and user comments together. We carefully labeled how the text and meme work together, what the poster intended, how the meme is structured, and how the community responded. Our tests with leading LVLMs show a clear weakness: models either fail to interpret critical information in the contexts, or overly focus on visual details while overlooking communicative purpose. MemeReaCon thus serves both as a diagnostic tool exposing current limitations and as a challenging benchmark to drive development toward more sophisticated LVLMs of the context-aware understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。