arXiv:2604.05930cs.CLcs.AI2026-04ACL被引 2

测试大模型能否看懂图文双关,发现多数模型难以分辨真伪双关。

"I See What You Did There": Can Large Vision-Language Models Understand Multimodal Puns?

  • 构建图文双关生成流程,设计包含干扰项的多类型双关数据集
  • 多数视觉语言模型在双关识别上表现差,准确率不足基准水平
  • 提出提示与模型改进策略,平均提升16.5%的识别准确率

双关是利用词义多重性与语音相似性制造幽默的修辞手法。在多模态双关中,视觉与文本元素协同作用,同时呈现字面意义与隐喻含义。尽管视觉语言模型(VLMs)广泛应用于多模态理解与生成,但其对双关的理解能力尚未系统评估,主要受限于缺乏严谨的评测基准。为此,我们首先提出一种多模态双关生成流程,并构建了MultiPun数据集,涵盖多种类型的双关及其对抗性非双关干扰项。评估显示,大多数模型难以区分真实双关与干扰项。此外,我们提出提示级和模型级改进策略,使平均F1得分提升16.5%。研究结果为未来具备跨模态推理能力、理解人类幽默的VLM发展提供了重要启示。

原文摘要 · Abstract (English)

Puns are a common form of rhetorical wordplay that exploits polysemy and phonetic similarity to create humor. In multimodal puns, visual and textual elements synergize to ground the literal sense and evoke the figurative meaning simultaneously. Although Vision-Language Models (VLMs) are widely used in multimodal understanding and generation, their ability to understand puns has not been systematically studied due to a scarcity of rigorous benchmarks. To address this, we first propose a multimodal pun generation pipeline. We then introduce MultiPun, a dataset comprising diverse types of puns alongside adversarial non-pun distractors. Our evaluation reveals that most models struggle to distinguish genuine puns from these distractors. Moreover, we propose both prompt-level and model-level strategies to enhance pun comprehension, with an average improvement of 16.5% in F1 scores. Our findings provide valuable insights for developing future VLMs that master the subtleties of human-like humor via cross-modal reasoning.

多模态理解视觉语言模型双关语幽默识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。