arXiv:2607.27798cs.AI2026-07被引 1

评测大模型解读文化梗的短板,发现知识缺失是核心问题

MemeBench: What LVLMs Miss When Interpreting Culture-Dependent Memes

论文配图:MemeBench: What LVLMs Miss When Interpreting Culture-Dependent Memes
图 1 · 摘自论文原文
  • 用分解式标注框架分析梗图中的视觉线索、身份关联、知识单元和推理机制
  • 26个模型普遍在知识理解上落后于视觉描述,最强模型仍有22.6%差距
  • 提出针对性检索方法KAR,能有效补足知识缺口且减少误判

大型视觉语言模型在描述视觉内容方面已有提升,但准确描述并不等于正确解读——当意义依赖像素之外的文化知识时,模型表现明显不足。梗图正是暴露这一差距的理想载体,因其依赖特定文化背景、社群惯例与隐性知识。现有梗图评测多简化为标签或整体评分,难以定位失败原因。我们提出MemeBench,一个包含1,253张中英文梗图的诊断型基准,覆盖动漫、漫画、游戏及周边亚文化,配有真人撰写参考答案与质量控制的VIKR标注。VIKR框架将解释拆解为视觉线索、身份关联、知识单元与推理机制四部分。在26个LVLM测试中,所有模型对可见内容的捕捉优于知识理解,最强模型仍存在22.6%的视觉-知识差距。为进一步验证诊断能力,我们引入KAR——基于CultureBase的实体引导检索基线。在四个受控模型中,KAR使VIKR成功率提升3.6%-7.4%,相比通用检索更有效修复错误答案并减少新错误。然而,两种检索方案均在提升身份与知识理解的同时,略微降低视觉覆盖。MemeBench不仅能判断解释是否成功,还能揭示缺失环节,并检验针对性证据能否填补诊断出的空缺。

原文摘要 · Abstract (English)

Large vision-language models have improved at describing visual content, but accurate descriptions do not ensure interpretation when meaning depends on knowledge beyond the pixels. Memes expose this gap because they rely on cultural entities, background knowledge, and community conventions. Most meme benchmarks reduce interpretation to labels or holistic scores, obscuring where an explanation breaks down. We introduce MemeBench, a diagnostic benchmark of 1,253 Chinese and English memes with human-written references and quality-controlled VIKR annotations, centered on anime, comics, games, and adjacent online subcultures. Its VIKR schema decomposes explanations into Visual clues, Identity links, Knowledge units, and Reasoning mechanisms. Across 26 LVLMs, every model covers visible content more reliably than the knowledge needed to interpret it, and even the strongest retains a 22.6% Visual-Knowledge gap. To test whether this diagnosis can guide improvement, we introduce KAR, an entity-guided retrieval baseline built on CultureBase. Across four controlled models, KAR raises VIKR Success by 3.6-7.4% and, compared with generic retrieval, repairs more answers and breaks fewer. Yet both retrieval conditions improve Identity and Knowledge while reducing Visual coverage in every comparison. MemeBench reveals whether an interpretation succeeds, what is missing, and whether targeted evidence fills the diagnosed gap.

视觉语言模型文化理解梗图评测知识缺口

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。