评测多模态模型对南亚文化梗的理解能力,发现加一点背景信息能显著提升表现。
MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal Models

- 构建1000个南亚多语言文化梗数据集,含上下文注释和人工解释
- 提供背景信息后,模型相似度提升11.8点,评分提高0.86分
- 开源数据集与代码,适合研究文化理解与多模态推理的学者
meme理解不仅需识别视觉内容或字面文本,更依赖隐含的文化知识与语用推理,而大多数视觉-语言模型仍缺乏此能力。本文提出MemeCULT-1K,一个包含1000个南亚文化梗的多语言基准,涵盖孟加拉语、英语和印地语,每个梗配以文化背景说明及三份人工撰写解释,并附带54个孟加拉地方方言梗。在两种设置下评估了13个主流视觉-语言模型(VLMs):仅看meme和带上下文。提供少量文化背景信息后,所有模型在各语言中均取得稳定提升:平均SBERT相似度从44.6升至56.4(+11.8),BLEURT从37.3升至42.3(+5.0),LLM-as-a-Judge得分从2.57升至3.43(满分5,+0.86)。细粒度错误分析显示,闭源模型主要问题在于实体与引用误认,而开源模型则受限于更广泛的文化知识缺失,语言与语音错误在各类模型中均最难以通过上下文修复。结果凸显文化驱动的meme理解难度,推动未来显式文化知识整合的研究。数据集与代码已公开于TawsifDipto17/MemeCULT-1K。
原文摘要 · Abstract (English)
Meme understanding goes beyond recognizing visual content or literal text; it requires implicit cultural knowledge and pragmatic inference that most vision-language models still lack. We introduce MemeCULT-1K, a multilingual benchmark of 1,000 South Asian memes in Bengali, English, and Hindi, where each meme is paired with a cultural context note and three human-written explanations, along with a supplementary set of 54 Bengali regional dialect memes. We evaluate thirteen popular Vision Language Models (VLMs) under two settings: meme-only and context-aware. Providing minimal cultural context yields consistent gains across all models and languages: mean SBERT similarity improves from 44.6 to 56.4 (+11.8), BLEURT from 37.3 to 42.3 (+5.0), and LLM-as-a-Judge scores from 2.57 to 3.43 out of 5 (+0.86). Fine-grained error analysis reveals that closed-source models fail mainly on entity and reference misidentification, while open-source models are bottlenecked by broader cultural knowledge gaps, with linguistic and phonological failures proving the most context-resistant across both. These results highlight the difficulty of culturally grounded meme understanding and motivate future work on explicit cultural knowledge integration. Our dataset and code are publicly available at TawsifDipto17/MemeCULT-1K.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。