arXiv:2603.23229cs.CL2026-03被引 2

评测大模型理解网络梗图隐喻意义的能力,发现普遍存在过度解读倾向。

I Came, I Saw, I Explained: Benchmarking Multimodal LLMs on Figurative Meaning in Memes

  • 选取8个主流多模态大模型,在3个数据集上测试其识别6类隐喻能力。
  • 所有模型均表现出强倾向性,即使无隐喻也误判为有,且解释常不忠实原图。
  • 适合关注AI对文化语境理解局限的研究者与内容安全开发者。

网络梗图是流行的多模态在线表达形式,常通过图文结合传递深层隐喻意义。然而,多模态大语言模型(MLLMs)如何融合并解析图像与文本以识别梗图中的隐喻仍不清晰。为此,我们在三个数据集上评估了八种前沿生成式多模态大模型在检测和解释六类隐喻意义方面的能力。同时,我们进行了人工评估,判断模型生成的解释是否支持预测标签,并是否忠实于原始梗图内容。结果表明,所有模型均表现出强烈的将梗图关联到隐喻意义的倾向,即使不存在此类意义。定性分析还显示,正确预测并不总伴随忠实的解释。

原文摘要 · Abstract (English)

Internet memes represent a popular form of multimodal online communication and often use figurative elements to convey layered meaning through the combination of text and images. However, it remains largely unclear how multimodal large language models (MLLMs) combine and interpret visual and textual information to identify figurative meaning in memes. To address this gap, we evaluate eight state-of-the-art generative MLLMs across three datasets on their ability to detect and explain six types of figurative meaning. In addition, we conduct a human evaluation of the explanations generated by these MLLMs, assessing whether the provided reasoning supports the predicted label and whether it remains faithful to the original meme content. Our findings indicate that all models exhibit a strong bias to associate a meme with figurative meaning, even when no such meaning is present. Qualitative analysis further shows that correct predictions are not always accompanied by faithful explanations.

多模态隐喻理解大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。